Patrick S. McCormick

dblp:147/4103 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0002-8305-6709ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Parallel and multicore computing · 68% Performance modeling and evaluation · 11% GPUs and heterogeneous computing · 6%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 50% Rendering · 50%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming runtimes
0.622021
Scaling implicit parallelism via dynamic control replication · PPoPP 2021
Task bench: a parameterized benchmark for evaluating parallel runtime performance · SC 2020
Parallel and multicore computing › parallelization strategies
implicit parallelism
0.622021
Scaling implicit parallelism via dynamic control replication · PPoPP 2021
Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017
Machine learning › Efficient and distributed learning
distributed training
0.612022
Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization · OSDI 2022
Parallel and multicore computing › parallel computing › parallel program analysis
dynamic dependence analysis
0.512021
Scaling implicit parallelism via dynamic control replication · PPoPP 2021
Parallel and multicore computing › parallel programming models
task-based programming
0.512021
Index launches: scalable, flexible representation of parallel task groups · SC 2021
Parallel and multicore computing › parallel programming runtimes
task-based runtime
0.512021
Index launches: scalable, flexible representation of parallel task groups · SC 2021
Performance modeling and evaluation
benchmarking
0.412020
Task bench: a parameterized benchmark for evaluating parallel runtime performance · SC 2020
Compilers and program optimization
parallelizing compiler
0.312017
Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017
Distributed systems
distributed graph processing
0.312017
A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017
Parallel and multicore computing
graph processing
0.312017
A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017
GPUs and heterogeneous computing › GPU graph processing
multi-GPU graph processing
0.312017
A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017
High-performance computing › scientific data analysis
in-situ analysis
0.212013
Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013
Energy-efficient computing
power modeling
0.212013
Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013
Visualization and visual analytics
flow visualization
0.112011
Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental Techniques · IEEE Trans. Vis. Comput. Graph. 2011
Performance modeling and evaluation › performance prediction
execution time prediction
0.112017
A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017
Parallel and multicore computing
parallel programming models
0.112017
Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017
Energy-efficient computing
power-performance tradeoff
0.012013
Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013
Computational science and engineering
computational fluid dynamics
0.012011
Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental Techniques · IEEE Trans. Vis. Comput. Graph. 2011

Methods — techniques the papers use, named apart from their topics

static program analysis · 1.0dynamic program analysis · 1.0logical regions · 0.6control replication · 0.6dynamic control replication · 0.5parameterized benchmarking · 0.4minimum effective task granularity · 0.4performance modeling · 0.3dynamic load balancing · 0.3schlieren imaging · 0.2interferometry · 0.2GPGPU · 0.2empirical power modeling · 0.2
YearPublicationVenuePosition
2022 Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization
Colin Unger, Wei Wu 0016, Sina Lin, Mandeep Baines, Carlos Efrain Quintero Narvaez, Vinay Ramakrishnaiah, Nirmal Prajapati, Patrick S. McCormick, Jamaludin Mohd-Yusof, Dheevatsa Mudigere, Jongsoo Park, Mikhail Smelyanskiy, Alex Aiken
OSDI9
2021 Scaling implicit parallelism via dynamic control replication
abstract
We present dynamic control replication, a run-time program analysis that enables scalable execution of implicitly parallel programs on large machines through a distributed and efficient dynamic dependence analysis. Dynamic control replication distributes dependence analysis by executing multiple copies of an implicitly parallel program while ensuring that they still collectively behave as a single execution. By distributing and parallelizing the dependence analysis, dynamic control replication supports efficient, on-the-fly computation of dependences for programs with arbitrary control flow at scale. We describe an asymptotically scalable algorithm for implementing dynamic control replication that maintains the sequential semantics of implicitly parallel programs.
Michael Bauer 0001, Wonchan Lee, Elliott Slaughter, Mario Di Renzo, Manolis Papadakis, Galen M. Shipman, Patrick S. McCormick, Michael Garland, Alex Aiken
PPoPP8
2021 Index launches: scalable, flexible representation of parallel task groups
abstract
It's common to see specialized language constructs in modern task-based programming systems for reasoning about groups of independent tasks intended for parallel execution. However, most systems use an ad-hoc representation that limits expressiveness and often overfits for a given application domain. We introduce index launches, a scalable and flexible representation of a group of tasks. Index launches use a flexible mechanism to indicate the data required for a given task, allowing them to be used for a much broader set of use cases while maintaining an efficient representation. We present a hybrid design for index launches, involving static and dynamic program analyses, along with a characterization of how they're used in Legion and Regent, and show how they generalize constructs found in other task-based systems. Finally, we present results of scaling experiments which demonstrate that index launches are crucial for the efficient distributed execution of several scientific codes in Regent.
Rupanshu Soi, Michael Bauer 0001, Sean Treichler, Manolis Papadakis, Wonchan Lee, Patrick S. McCormick, Alex Aiken, Elliott Slaughter
SC6
2020 Task bench: a parameterized benchmark for evaluating parallel runtime performance
abstract
We present Task Bench, a parameterized benchmark designed to explore the performance of distributed programming systems under a variety of application scenarios. Task Bench dramatically lowers the barrier to benchmarking and comparing multiple programming systems by making the implementation for a given system orthogonal to the benchmarks themselves: every benchmark constructed with Task Bench runs on every Task Bench implementation. Furthermore, Task Bench's parameterization enables a wide variety of benchmark scenarios that distill the key characteristics of larger applications. To assess the effectiveness and overheads of the tested systems, we introduce a novel metric, minimum effective task granularity (METG). We conduct a comprehensive study with 15 programming systems on up to 256 Haswell nodes of the Cori supercomputer. Running at scale, 100μs-long tasks are the finest granularity that any system runs efficiently with current technologies. We also study each system's scalability, ability to hide communication and mitigate load imbalance.
Elliott Slaughter, Wei Wu 0016, Yuankun Fu, Legend Brandenburg, Nicolai Garcia, Wilhem Kautz, Emily Marx, Kaleb S. Morris, Qinglei Cao, George Bosilca, Seema Mirchandaney, Wonchan Lee, Sean Treichler, Patrick S. McCormick, Alex Aiken
SC14
2020 On the memory attribution problem: A solution and case study using MPI
abstract
Summary As parallel applications running on large‐scale computing systems become increasingly memory constrained, the ability to attribute memory usage to the various components of the application is becoming increasingly important. We present the design and implementation of memnesia, a novel memory usage profiler for parallel and distributed message‐passing applications. Our approach captures both application– and message‐passing library–specific memory usage statistics from unmodified binaries dynamically linked to a message‐passing communication library. Using microbenchmarks and proxy applications, we evaluated our profiler across three Message Passing Interface (MPI) implementations and two hardware platforms. The results show that our approach and the corresponding implementation can accurately quantify memory resource usage as a function of time, scale, communication workload, and software or hardware system architecture, clearly distinguishing between application and MPI library memory usage at a per‐process level. With this new capability, we show that job size, communication workload, and hardware/software architecture influence peak runtime memory usage. In practice, this tool provides a potentially valuable source of information for application developers seeking to measure and optimize memory usage.
Samuel K. Gutierrez, Dorian C. Arnold, Kei Davis, Patrick S. McCormick
Concurr. Comput. Pract. Exp.4
2018 Isometry: A Path-Based Distributed Data Transfer System
abstract
Data transfers in parallel systems have a significant impact on the performance of applications. Most existing systems generally support only data transfers between memories with a direct hardware connection and have limited facilities for handling transformations to the data's layout in memory. As a result, to move data between memories that are not directly connected, higher levels of the software stack must explicitly divide a multi-hop transfer into a sequence of single-hop transfers and decide how and where to perform data layout conversions if needed. This approach results in inefficiencies, as the higher levels lack enough information to plan transfers as a whole, while the lower level that does the transfer sees only the individual single-hop requests.
Sean Treichler, Galen M. Shipman, Patrick S. McCormick, Alex Aiken
ICS4
2017 Integrating External Resources with a Task-Based Programming Model
abstract
Accessing external resources (e.g., loading input data, checkpointing snapshots, and out-of-core processing) can have a significant impact on the performance of supercomputer applications. However, no existing programming systems for high-performance computing directly manage and optimize these external accesses. As a result, users must explicitly manage external accesses alongside their computation at the application level, which can result in both correctness and performance issues. We address this limitation by introducing Iris, a task-based programming model with semantics for external resources. Iris allows applications to describe their access requirements to external resources and the relationship of those accesses to the computation. Iris incorporates external I/O into a deferred execution model, reschedules external I/O to overlap I/O with computation, and reduces external I/O when possible. We evaluate Iris on three microbenchmarks representative of important workloads in HPC and a full combustion simulation, S3D. We demonstrate that the Iris implementation of S3D reduces the external I/O overhead by up to 20×, compared to the Legion and the Fortran implementations.
Sean Treichler, Galen M. Shipman, Michael Bauer 0001, Noah Watkins, Carlos Maltzahn, Patrick S. McCormick, Alex Aiken
HiPC7
2017 Accommodating Thread-Level Heterogeneity in Coupled Parallel Applications
abstract
Hybrid parallel program models that combine message passing and multithreading (MP+MT) are becoming more popular, extending the basic message passing (MP) model that uses single-threaded processes for both inter- and intra-node parallelism. A consequence is that coupled parallel applications increasingly comprise MP libraries together with MP+MT libraries with differing preferred degrees of threading, resulting in thread-level heterogeneity. Retroactively matching threading levels between independently developed and maintained libraries is difficult; the challenge is exacerbated because contemporary parallel job launchers provide only static resource binding policies over entire application executions. A standard approach for accommodating thread-level heterogeneity is to under-subscribe compute resources such that the library with the highest degree of threading per process has one processing element per thread. This results in libraries with fewer threads per process utilizing only a fraction of the available compute resources. We present and evaluate a novel approach for accommodating thread-level heterogeneity. Our approach enables full utilization of all available compute resources throughout an application's execution by providing programmable facilities to dynamically reconfigure runtime environments for compute phases with differing threading factors and memory affinities. We show that our approach can improve overall application performance by up to 5.8× in real-world production codes. Furthermore, the practicality and utility of our approach has been demonstrated by continuous production use for over one year, and by more recent incorporation into a number of production codes.
Samuel K. Gutierrez, Kei Davis, Dorian C. Arnold, Randal S. Baker, Robert W. Robey, Patrick S. McCormick, Daniel Holladay, Jon A. Dahl, Joe Zerr, Florian Weik, Christoph Junghans
IPDPS6
2017 Control replication: compiling implicit parallelism to efficient SPMD with logical regions
abstract
We present control replication, a technique for generating high-performance and scalable SPMD code from implicitly parallel programs. In contrast to traditional parallel programming models that require the programmer to explicitly manage threads and the communication and synchronization between them, implicitly parallel programs have sequential execution semantics and naturally avoid the pitfalls of explicitly parallel code. However, without optimizations to distribute control overhead, scalability is often poor.
Elliott Slaughter, Wonchan Lee, Sean Treichler, Michael Bauer 0001, Galen M. Shipman, Patrick S. McCormick, Alex Aiken
SC7
2017 A Distributed Multi-GPU System for Fast Graph Processing
abstract
We present Lux, a distributed multi-GPU system that achieves fast graph processing by exploiting the aggregate memory bandwidth of multiple GPUs and taking advantage of locality in the memory hierarchy of multi-GPU clusters. Lux provides two execution models that optimize algorithmic efficiency and enable important GPU optimizations, respectively. Lux also uses a novel dynamic load balancing strategy that is cheap and achieves good load balance across GPUs. In addition, we present a performance model that quantitatively predicts the execution times and automatically selects the runtime configurations for Lux applications. Experiments show that Lux achieves up to 20X speedup over state-of-the-art shared memory systems and up to two orders of magnitude speedup over distributed systems.
Yongkee Kwon, Galen M. Shipman, Patrick S. McCormick, Mattan Erez, Alex Aiken
Proc. VLDB Endow.4
2013 Exploring power behaviors and trade-offs of in-situ data analytics
abstract
As scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems.
Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky
SC10
2012 Automatic NUMA characterization using Cbench
abstract
Clusters of seemingly homogeneous compute nodes are increasingly heterogeneous within each node due to replication and distribution of node-level subsystems. This intra-node heterogeneity can adversely affect program execution performance by inflicting additional data-access costs when accessing non-local data. In this work-in-progress paper, we present extensions to the Cbench Scalable Testing Framework for analyzing main memory and PCIe data-access performance in modern NUMA architectures. The information provided by this tool will be of use for task scheduling, performance modeling, and evaluation of NUMA systems.
Ryan K. Braithwaite, Wu-chun Feng, Patrick S. McCormick
ICPE3
2011 Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental Techniques
abstract
Understanding fluid flow is a difficult problem and of increasing importance as computational fluid dynamics (CFD) produces an abundance of simulation data. Experimental flow analysis has employed techniques such as shadowgraph, interferometry, and schlieren imaging for centuries, which allow empirical observation of inhomogeneous flows. Shadowgraphs provide an intuitive way of looking at small changes in flow dynamics through caustic effects while schlieren cutoffs introduce an intensity gradation for observing large scale directional changes in the flow. Interferometry tracks changes in phase-shift resulting in bands appearing. The combination of these shading effects provides an informative global analysis of overall fluid flow. Computational solutions for these methods have proven too complex until recently due to the fundamental physical interaction of light refracting through the flow field. In this paper, we introduce a novel method to simulate the refraction of light to generate synthetic shadowgraph, schlieren and interferometry images of time-varying scalar fields derived from computational fluid dynamics data. Our method computes physically accurate schlieren and shadowgraph images at interactive rates by utilizing a combination of GPGPU programming, acceleration methods, and data-dependent probabilistic schlieren cutoffs. Applications of our method to multifield data and custom application-dependent color filter creation are explored. Results comparing this method to previous schlieren approximations are finally presented.
Carson Brownlee, Vincent Pegoraro, Siddharth Shankar, Patrick S. McCormick, Charles D. Hansen
IEEE Trans. Vis. Comput. Graph.4
2010 Physically-based interactive schlieren flow visualization
abstract
Understanding fluid flow is a difficult problem and of increasing importance as computational fluid dynamics produces an abundance of simulation data. Experimental flow analysis has employed techniques such as shadowgraph and schlieren imaging for centuries which allow empirical observation of inhomogeneous flows. Shadowgraphs provide an intuitive way of looking at small changes in flow dynamics through caustic effects while schlieren cutoffs introduce an intensity gradation for observing large scale directional changes in the flow. The combination of these shading effects provides an informative global analysis of overall fluid flow. Computational solutions for these methods have proven too complex until recently due to the fundamental physical interaction of light refracting through the flow field. In this paper, we introduce a novel method to simulate the refraction of light to generate synthetic shadowgraphs and schlieren images of time-varying scalar fields derived from computational fluid dynamics (CFD) data. Our method computes physically accurate schlieren and shadowgraph images at interactive rates by utilizing a combination of GPGPU programming, acceleration methods, and data-dependent probabilistic schlieren cutoffs. Results comparing this method to previous schlieren approximations are presented.
Carson Brownlee, Vincent Pegoraro, Singer Shankar, Patrick S. McCormick, Charles D. Hansen
PacificVis4
2007 Exploring weak scalability for FEM calculations on a GPU-enhanced cluster
Dominik Göddeke, Robert Strzodka, Jamaludin Mohd-Yusof, Patrick S. McCormick, Sven H. M. Buijssen, Matthias Grajewski, Stefan Turek
Parallel Comput.4
2007 Scout: a data-parallel programming language for graphics processors
Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Jamaludin Mohd-Yusof, Greg Roth, Sharen J. Cummins
Parallel Comput.1
2005 General Purpose Computation on Graphics Hardware
Aaron E. Lefohn, Ian Buck, Patrick S. McCormick, John D. Owens, Timothy J. Purcell, Robert Strzodka
IEEE Visualization3
2004 Scout: A Hardware-Accelerated System for Quantitatively Driven Visualization and Analysis
abstract
Quantitative techniques for visualization are critical to the successful analysis of both acquired and simulated scientific data. Many visualization techniques rely on indirect mappings, such as transfer functions, to produce the final imagery. In many situations, it is preferable and more powerful to express these mappings as mathematical expressions, or queries, that can then be directly applied to the data. We present a hardware-accelerated system that provides such capabilities and exploits current graphics hardware for portions of the computational tasks that would otherwise be executed on the CPU. In our approach, the direct programming of the graphics processor using a concise data parallel language, gives scientists the capability to efficiently explore and visualize data sets.
Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Charles D. Hansen, Greg Roth
IEEE Visualization1
2003 Visualizing Industrial CT Volume Data for Nondestructive Testing Applications
abstract
This paper describes a set of techniques developed for the visualization of high-resolution volume data generated from industrial computed tomography for nondestructive testing (NDT) applications. Because the data are typically noisy and contain fine features, direct volume rendering methods do not always give us satisfactory results. We have coupled region growing techniques and a 2D histogram interface to facilitate volumetric feature extraction. The new interface allows the user to conveniently identify, separate or composite, and compare features in the data. To lower the cost of segmentation, we show how partial region growing results can suggest a reasonably good classification function for the rendering of the whole volume. The NDT applications that we work on demand visualization tasks including not only feature extraction and visual inspection, but also modeling and measurement of concealed structures in volumetric objects. An efficient filtering and modeling process for generating surface representation of extracted features is also introduced. Four CT data sets for preliminary NDT are used to demonstrate the effectiveness of the new visualization strategy that we have developed.
Runzhen Huang, Kwan-Liu Ma, Patrick S. McCormick, William Ward
IEEE Visualization3
1997 Wildfire visualization (case study)
abstract
The ability to forecast the progress of crisis events would significantly reduce human suffering and loss of life, the destruction of property and expenditures for assessment and recovery. Los Alamos National Laboratory has established a scientific thrust in crisis forecasting to address this national challenge. In the initial phase of this project, scientists at Los Alamos are developing computer models to predict the spread of a wildfire. Visualization of the results of the wildfire simulation will be used by scientists to assess the quality of the simulation and eventually by fire personnel as a visual forecast of the wildfire's evolution. The fire personnel and scientists want the visualization to look as realistic as possible without compromising scientific accuracy. This paper describes how the visualization was created, analyzes the tools and approach that were used, and suggests directions for future work and research.
James P. Ahrens, Patrick S. McCormick, James Bossert, Jon Reisner, Judith Winterkamp
IEEE Visualization2