EDBT 2026 Demo / reviewers in the wild / expert
Patrick S. McCormick
dblp:147/4103
· DBLP profile ↗
20ranked-venue papers
2as first author
3since 2021 · last 2022
0000-0002-8305-6709ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Parallel and multicore computing · 68% Performance modeling and evaluation · 11% GPUs and heterogeneous computing · 6% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 100% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 50% Rendering · 50% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing
parallel programming runtimes |
0.6 | 2 | 2021 | Scaling implicit parallelism via dynamic control replication · PPoPP 2021 Task bench: a parameterized benchmark for evaluating parallel runtime performance · SC 2020 |
Parallel and multicore computing › parallelization strategies
implicit parallelism |
0.6 | 2 | 2021 | Scaling implicit parallelism via dynamic control replication · PPoPP 2021 Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017 |
Machine learning › Efficient and distributed learning
distributed training |
0.6 | 1 | 2022 | Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization · OSDI 2022 |
Parallel and multicore computing › parallel computing › parallel program analysis
dynamic dependence analysis |
0.5 | 1 | 2021 | Scaling implicit parallelism via dynamic control replication · PPoPP 2021 |
Parallel and multicore computing › parallel programming models
task-based programming |
0.5 | 1 | 2021 | Index launches: scalable, flexible representation of parallel task groups · SC 2021 |
Parallel and multicore computing › parallel programming runtimes
task-based runtime |
0.5 | 1 | 2021 | Index launches: scalable, flexible representation of parallel task groups · SC 2021 |
Performance modeling and evaluation
benchmarking |
0.4 | 1 | 2020 | Task bench: a parameterized benchmark for evaluating parallel runtime performance · SC 2020 |
Compilers and program optimization
parallelizing compiler |
0.3 | 1 | 2017 | Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017 |
Distributed systems
distributed graph processing |
0.3 | 1 | 2017 | A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017 |
Parallel and multicore computing
graph processing |
0.3 | 1 | 2017 | A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017 |
GPUs and heterogeneous computing › GPU graph processing
multi-GPU graph processing |
0.3 | 1 | 2017 | A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017 |
High-performance computing › scientific data analysis
in-situ analysis |
0.2 | 1 | 2013 | Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013 |
Energy-efficient computing
power modeling |
0.2 | 1 | 2013 | Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013 |
Visualization and visual analytics
flow visualization |
0.1 | 1 | 2011 | Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental Techniques · IEEE Trans. Vis. Comput. Graph. 2011 |
Performance modeling and evaluation › performance prediction
execution time prediction |
0.1 | 1 | 2017 | A Distributed Multi-GPU System for Fast Graph Processing · Proc. VLDB Endow. 2017 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2017 | Control replication: compiling implicit parallelism to efficient SPMD with logical regions · SC 2017 |
Energy-efficient computing
power-performance tradeoff |
0.0 | 1 | 2013 | Exploring power behaviors and trade-offs of in-situ data analytics · SC 2013 |
Computational science and engineering
computational fluid dynamics |
0.0 | 1 | 2011 | Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental Techniques · IEEE Trans. Vis. Comput. Graph. 2011 |
Methods — techniques the papers use, named apart from their topics
static program analysis · 1.0dynamic program analysis · 1.0logical regions · 0.6control replication · 0.6dynamic control replication · 0.5parameterized benchmarking · 0.4minimum effective task granularity · 0.4performance modeling · 0.3dynamic load balancing · 0.3schlieren imaging · 0.2interferometry · 0.2GPGPU · 0.2empirical power modeling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Unity: Accelerating DNN Training Through Joint Optimization of Algebraic Transformations and Parallelization
Colin Unger, Wei Wu 0016, Sina Lin, Mandeep Baines, Carlos Efrain Quintero Narvaez, Vinay Ramakrishnaiah, Nirmal Prajapati, Patrick S. McCormick, Jamaludin Mohd-Yusof, Dheevatsa Mudigere, Jongsoo Park, Mikhail Smelyanskiy, Alex Aiken |
OSDI | 9 |
| 2021 | Scaling implicit parallelism via dynamic control replicationabstractWe present dynamic control replication, a run-time program analysis that enables scalable execution of implicitly parallel programs on large machines through a distributed and efficient dynamic dependence analysis. Dynamic control replication distributes dependence analysis by executing multiple copies of an implicitly parallel program while ensuring that they still collectively behave as a single execution. By distributing and parallelizing the dependence analysis, dynamic control replication supports efficient, on-the-fly computation of dependences for programs with arbitrary control flow at scale. We describe an asymptotically scalable algorithm for implementing dynamic control replication that maintains the sequential semantics of implicitly parallel programs. Michael Bauer 0001, Wonchan Lee, Elliott Slaughter, Mario Di Renzo, Manolis Papadakis, Galen M. Shipman, Patrick S. McCormick, Michael Garland, Alex Aiken |
PPoPP | 8 |
| 2021 | Index launches: scalable, flexible representation of parallel task groupsabstractIt's common to see specialized language constructs in modern task-based programming systems for reasoning about groups of independent tasks intended for parallel execution. However, most systems use an ad-hoc representation that limits expressiveness and often overfits for a given application domain. We introduce index launches, a scalable and flexible representation of a group of tasks. Index launches use a flexible mechanism to indicate the data required for a given task, allowing them to be used for a much broader set of use cases while maintaining an efficient representation. We present a hybrid design for index launches, involving static and dynamic program analyses, along with a characterization of how they're used in Legion and Regent, and show how they generalize constructs found in other task-based systems. Finally, we present results of scaling experiments which demonstrate that index launches are crucial for the efficient distributed execution of several scientific codes in Regent. Rupanshu Soi, Michael Bauer 0001, Sean Treichler, Manolis Papadakis, Wonchan Lee, Patrick S. McCormick, Alex Aiken, Elliott Slaughter |
SC | 6 |
| 2020 | Task bench: a parameterized benchmark for evaluating parallel runtime performanceabstractWe present Task Bench, a parameterized benchmark designed to explore the performance of distributed programming systems under a variety of application scenarios. Task Bench dramatically lowers the barrier to benchmarking and comparing multiple programming systems by making the implementation for a given system orthogonal to the benchmarks themselves: every benchmark constructed with Task Bench runs on every Task Bench implementation. Furthermore, Task Bench's parameterization enables a wide variety of benchmark scenarios that distill the key characteristics of larger applications. To assess the effectiveness and overheads of the tested systems, we introduce a novel metric, minimum effective task granularity (METG). We conduct a comprehensive study with 15 programming systems on up to 256 Haswell nodes of the Cori supercomputer. Running at scale, 100μs-long tasks are the finest granularity that any system runs efficiently with current technologies. We also study each system's scalability, ability to hide communication and mitigate load imbalance. Elliott Slaughter, Wei Wu 0016, Yuankun Fu, Legend Brandenburg, Nicolai Garcia, Wilhem Kautz, Emily Marx, Kaleb S. Morris, Qinglei Cao, George Bosilca, Seema Mirchandaney, Wonchan Lee, Sean Treichler, Patrick S. McCormick, Alex Aiken |
SC | 14 |
| 2020 | On the memory attribution problem: A solution and case study using MPIabstractSummary As parallel applications running on large‐scale computing systems become increasingly memory constrained, the ability to attribute memory usage to the various components of the application is becoming increasingly important. We present the design and implementation of memnesia, a novel memory usage profiler for parallel and distributed message‐passing applications. Our approach captures both application– and message‐passing library–specific memory usage statistics from unmodified binaries dynamically linked to a message‐passing communication library. Using microbenchmarks and proxy applications, we evaluated our profiler across three Message Passing Interface (MPI) implementations and two hardware platforms. The results show that our approach and the corresponding implementation can accurately quantify memory resource usage as a function of time, scale, communication workload, and software or hardware system architecture, clearly distinguishing between application and MPI library memory usage at a per‐process level. With this new capability, we show that job size, communication workload, and hardware/software architecture influence peak runtime memory usage. In practice, this tool provides a potentially valuable source of information for application developers seeking to measure and optimize memory usage. Samuel K. Gutierrez, Dorian C. Arnold, Kei Davis, Patrick S. McCormick |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Isometry: A Path-Based Distributed Data Transfer SystemabstractData transfers in parallel systems have a significant impact on the performance of applications. Most existing systems generally support only data transfers between memories with a direct hardware connection and have limited facilities for handling transformations to the data's layout in memory. As a result, to move data between memories that are not directly connected, higher levels of the software stack must explicitly divide a multi-hop transfer into a sequence of single-hop transfers and decide how and where to perform data layout conversions if needed. This approach results in inefficiencies, as the higher levels lack enough information to plan transfers as a whole, while the lower level that does the transfer sees only the individual single-hop requests. Sean Treichler, Galen M. Shipman, Patrick S. McCormick, Alex Aiken |
ICS | 4 |
| 2017 | Integrating External Resources with a Task-Based Programming ModelabstractAccessing external resources (e.g., loading input data, checkpointing snapshots, and out-of-core processing) can have a significant impact on the performance of supercomputer applications. However, no existing programming systems for high-performance computing directly manage and optimize these external accesses. As a result, users must explicitly manage external accesses alongside their computation at the application level, which can result in both correctness and performance issues. We address this limitation by introducing Iris, a task-based programming model with semantics for external resources. Iris allows applications to describe their access requirements to external resources and the relationship of those accesses to the computation. Iris incorporates external I/O into a deferred execution model, reschedules external I/O to overlap I/O with computation, and reduces external I/O when possible. We evaluate Iris on three microbenchmarks representative of important workloads in HPC and a full combustion simulation, S3D. We demonstrate that the Iris implementation of S3D reduces the external I/O overhead by up to 20×, compared to the Legion and the Fortran implementations. Sean Treichler, Galen M. Shipman, Michael Bauer 0001, Noah Watkins, Carlos Maltzahn, Patrick S. McCormick, Alex Aiken |
HiPC | 7 |
| 2017 | Accommodating Thread-Level Heterogeneity in Coupled Parallel ApplicationsabstractHybrid parallel program models that combine message passing and multithreading (MP+MT) are becoming more popular, extending the basic message passing (MP) model that uses single-threaded processes for both inter- and intra-node parallelism. A consequence is that coupled parallel applications increasingly comprise MP libraries together with MP+MT libraries with differing preferred degrees of threading, resulting in thread-level heterogeneity. Retroactively matching threading levels between independently developed and maintained libraries is difficult; the challenge is exacerbated because contemporary parallel job launchers provide only static resource binding policies over entire application executions. A standard approach for accommodating thread-level heterogeneity is to under-subscribe compute resources such that the library with the highest degree of threading per process has one processing element per thread. This results in libraries with fewer threads per process utilizing only a fraction of the available compute resources. We present and evaluate a novel approach for accommodating thread-level heterogeneity. Our approach enables full utilization of all available compute resources throughout an application's execution by providing programmable facilities to dynamically reconfigure runtime environments for compute phases with differing threading factors and memory affinities. We show that our approach can improve overall application performance by up to 5.8× in real-world production codes. Furthermore, the practicality and utility of our approach has been demonstrated by continuous production use for over one year, and by more recent incorporation into a number of production codes. Samuel K. Gutierrez, Kei Davis, Dorian C. Arnold, Randal S. Baker, Robert W. Robey, Patrick S. McCormick, Daniel Holladay, Jon A. Dahl, Joe Zerr, Florian Weik, Christoph Junghans |
IPDPS | 6 |
| 2017 | Control replication: compiling implicit parallelism to efficient SPMD with logical regionsabstractWe present control replication, a technique for generating high-performance and scalable SPMD code from implicitly parallel programs. In contrast to traditional parallel programming models that require the programmer to explicitly manage threads and the communication and synchronization between them, implicitly parallel programs have sequential execution semantics and naturally avoid the pitfalls of explicitly parallel code. However, without optimizations to distribute control overhead, scalability is often poor. Elliott Slaughter, Wonchan Lee, Sean Treichler, Michael Bauer 0001, Galen M. Shipman, Patrick S. McCormick, Alex Aiken |
SC | 7 |
| 2017 | A Distributed Multi-GPU System for Fast Graph ProcessingabstractWe present Lux, a distributed multi-GPU system that achieves fast graph processing by exploiting the aggregate memory bandwidth of multiple GPUs and taking advantage of locality in the memory hierarchy of multi-GPU clusters. Lux provides two execution models that optimize algorithmic efficiency and enable important GPU optimizations, respectively. Lux also uses a novel dynamic load balancing strategy that is cheap and achieves good load balance across GPUs. In addition, we present a performance model that quantitatively predicts the execution times and automatically selects the runtime configurations for Lux applications. Experiments show that Lux achieves up to 20X speedup over state-of-the-art shared memory systems and up to two orders of magnitude speedup over distributed systems. Yongkee Kwon, Galen M. Shipman, Patrick S. McCormick, Mattan Erez, Alex Aiken |
Proc. VLDB Endow. | 4 |
| 2013 | Exploring power behaviors and trade-offs of in-situ data analyticsabstractAs scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems. Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky |
SC | 10 |
| 2012 | Automatic NUMA characterization using CbenchabstractClusters of seemingly homogeneous compute nodes are increasingly heterogeneous within each node due to replication and distribution of node-level subsystems. This intra-node heterogeneity can adversely affect program execution performance by inflicting additional data-access costs when accessing non-local data. In this work-in-progress paper, we present extensions to the Cbench Scalable Testing Framework for analyzing main memory and PCIe data-access performance in modern NUMA architectures. The information provided by this tool will be of use for task scheduling, performance modeling, and evaluation of NUMA systems. Ryan K. Braithwaite, Wu-chun Feng, Patrick S. McCormick |
ICPE | 3 |
| 2011 | Physically-Based Interactive Flow Visualization Based on Schlieren and Interferometry Experimental TechniquesabstractUnderstanding fluid flow is a difficult problem and of increasing importance as computational fluid dynamics (CFD) produces an abundance of simulation data. Experimental flow analysis has employed techniques such as shadowgraph, interferometry, and schlieren imaging for centuries, which allow empirical observation of inhomogeneous flows. Shadowgraphs provide an intuitive way of looking at small changes in flow dynamics through caustic effects while schlieren cutoffs introduce an intensity gradation for observing large scale directional changes in the flow. Interferometry tracks changes in phase-shift resulting in bands appearing. The combination of these shading effects provides an informative global analysis of overall fluid flow. Computational solutions for these methods have proven too complex until recently due to the fundamental physical interaction of light refracting through the flow field. In this paper, we introduce a novel method to simulate the refraction of light to generate synthetic shadowgraph, schlieren and interferometry images of time-varying scalar fields derived from computational fluid dynamics data. Our method computes physically accurate schlieren and shadowgraph images at interactive rates by utilizing a combination of GPGPU programming, acceleration methods, and data-dependent probabilistic schlieren cutoffs. Applications of our method to multifield data and custom application-dependent color filter creation are explored. Results comparing this method to previous schlieren approximations are finally presented. Carson Brownlee, Vincent Pegoraro, Siddharth Shankar, Patrick S. McCormick, Charles D. Hansen |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2010 | Physically-based interactive schlieren flow visualizationabstractUnderstanding fluid flow is a difficult problem and of increasing importance as computational fluid dynamics produces an abundance of simulation data. Experimental flow analysis has employed techniques such as shadowgraph and schlieren imaging for centuries which allow empirical observation of inhomogeneous flows. Shadowgraphs provide an intuitive way of looking at small changes in flow dynamics through caustic effects while schlieren cutoffs introduce an intensity gradation for observing large scale directional changes in the flow. The combination of these shading effects provides an informative global analysis of overall fluid flow. Computational solutions for these methods have proven too complex until recently due to the fundamental physical interaction of light refracting through the flow field. In this paper, we introduce a novel method to simulate the refraction of light to generate synthetic shadowgraphs and schlieren images of time-varying scalar fields derived from computational fluid dynamics (CFD) data. Our method computes physically accurate schlieren and shadowgraph images at interactive rates by utilizing a combination of GPGPU programming, acceleration methods, and data-dependent probabilistic schlieren cutoffs. Results comparing this method to previous schlieren approximations are presented. Carson Brownlee, Vincent Pegoraro, Singer Shankar, Patrick S. McCormick, Charles D. Hansen |
PacificVis | 4 |
| 2007 | Exploring weak scalability for FEM calculations on a GPU-enhanced cluster
Dominik Göddeke, Robert Strzodka, Jamaludin Mohd-Yusof, Patrick S. McCormick, Sven H. M. Buijssen, Matthias Grajewski, Stefan Turek |
Parallel Comput. | 4 |
| 2007 | Scout: a data-parallel programming language for graphics processors
Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Jamaludin Mohd-Yusof, Greg Roth, Sharen J. Cummins |
Parallel Comput. | 1 |
| 2005 | General Purpose Computation on Graphics Hardware
Aaron E. Lefohn, Ian Buck, Patrick S. McCormick, John D. Owens, Timothy J. Purcell, Robert Strzodka |
IEEE Visualization | 3 |
| 2004 | Scout: A Hardware-Accelerated System for Quantitatively Driven Visualization and AnalysisabstractQuantitative techniques for visualization are critical to the successful analysis of both acquired and simulated scientific data. Many visualization techniques rely on indirect mappings, such as transfer functions, to produce the final imagery. In many situations, it is preferable and more powerful to express these mappings as mathematical expressions, or queries, that can then be directly applied to the data. We present a hardware-accelerated system that provides such capabilities and exploits current graphics hardware for portions of the computational tasks that would otherwise be executed on the CPU. In our approach, the direct programming of the graphics processor using a concise data parallel language, gives scientists the capability to efficiently explore and visualize data sets. Patrick S. McCormick, Jeff T. Inman, James P. Ahrens, Charles D. Hansen, Greg Roth |
IEEE Visualization | 1 |
| 2003 | Visualizing Industrial CT Volume Data for Nondestructive Testing ApplicationsabstractThis paper describes a set of techniques developed for the visualization of high-resolution volume data generated from industrial computed tomography for nondestructive testing (NDT) applications. Because the data are typically noisy and contain fine features, direct volume rendering methods do not always give us satisfactory results. We have coupled region growing techniques and a 2D histogram interface to facilitate volumetric feature extraction. The new interface allows the user to conveniently identify, separate or composite, and compare features in the data. To lower the cost of segmentation, we show how partial region growing results can suggest a reasonably good classification function for the rendering of the whole volume. The NDT applications that we work on demand visualization tasks including not only feature extraction and visual inspection, but also modeling and measurement of concealed structures in volumetric objects. An efficient filtering and modeling process for generating surface representation of extracted features is also introduced. Four CT data sets for preliminary NDT are used to demonstrate the effectiveness of the new visualization strategy that we have developed. Runzhen Huang, Kwan-Liu Ma, Patrick S. McCormick, William Ward |
IEEE Visualization | 3 |
| 1997 | Wildfire visualization (case study)abstractThe ability to forecast the progress of crisis events would significantly reduce human suffering and loss of life, the destruction of property and expenditures for assessment and recovery. Los Alamos National Laboratory has established a scientific thrust in crisis forecasting to address this national challenge. In the initial phase of this project, scientists at Los Alamos are developing computer models to predict the spread of a wildfire. Visualization of the results of the wildfire simulation will be used by scientists to assess the quality of the simulation and eventually by fire personnel as a visual forecast of the wildfire's evolution. The fire personnel and scientists want the visualization to look as realistic as possible without compromising scientific accuracy. This paper describes how the visualization was created, analyzes the tools and approach that were used, and suggests directions for future work and research. James P. Ahrens, Patrick S. McCormick, James Bossert, Jon Reisner, Judith Winterkamp |
IEEE Visualization | 2 |