VLDB 2026 Research / reviewers in the wild / expert
Nicholas J. Wright
dblp:20/7036
· DBLP profile ↗
26ranked-venue papers
0as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 12 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Energy-Efficient HPC: Insights from Power Profiling a Cloud-Resolving Earth System Model
Zhengji Zhao, Noel Keen, Oscar Antepara, Samuel Williams 0001, Luca Bertagna, Naser Mahfouz, James B. White III, Leonid Oliker, Brian Austin, Nicholas J. Wright |
IPDPS | 10 |
| 2025 | A Global Perspective on Supercomputer Power Provisioning: Case Studies from United States and EuropeabstractElectrical provisioning in high performance computing is transitioning from simple nameplate Thermal Design Power (TDP) models to more nuanced approaches based on expected electrical load.This paper captures current power Tapasya Patki, Barry Rountree, Torsten Wilde, Andrea Bartolini, Stephanie Brink, Esa Heiskanen, Sachin Idgunji, Matthias Maiterth, James H. Rogers, Ermal Rrapaj, Ralf Schneider, Woong Shin, Kathleen Shoga, Christian Simmendinger, Nicholas J. Wright, Zhengji Zhao |
ICS | 15 |
| 2025 | Benchmark-driven Models for Energy Analysis and Attribution of GPU-Accelerated SupercomputingabstractAs advances in energy-efficiency become the primary limiter to increases in power-constrained supercomputing and machine learning performance, it is imperative developers, architects, and practitioners understand how modern GPUs consume energy when running HPC and ML applications. Rather than opaque coarse-grained metrics, in this paper, we develop an extensible, microbenchmark-parameterized energy model capable of attributing application energy not only by functional unit (FPU, tensor core, integer ALU) and memory level (L1, L2, HBM), but can also differentiate control energy from datapath energy. We examine trends in energy per operation among four generations of GPUs and validate our results using supercomputing and ML/AI procurement workloads. Our insights and extrapolations can be used to drive the future of CMOS and memory technologies, computer architecture research, algorithmic innovation, optimizations for power-constrained and mobile environments, and data center operations. Oscar Antepara, Zhengji Zhao, Brian Austin, Nan Ding 0006, Leonid Oliker, Nicholas J. Wright, Samuel Williams 0001 |
SC | 6 |
| 2024 | A Workflow Roofline Model for End-to-End Workflow Performance AnalysisabstractAs next-generation experimental and observational instruments for scientific research are being deployed with higher resolutions and faster data capture rates, the fundamental demands of producing high-quality scientific throughput require portability and performance to meet the high productivity goals. Understanding such a workflow’s end-to-end performance on HPC systems is formidable work. In this paper, we address this challenge by introducing a Workflow Roofline model, which ties a workflow’s end-to-end performance with peak node- and system-performance constraints. We analyze four workflows: LCLS, a time-sensitive workflow that is bound by system external bandwidth; BerkeleyGW, a traditional HPC workflow that is bound by node-local performance; CosmoFlow, an AI workflow that is bound by the CPU preprocessing; and GPTune, an auto tuner that is bound by the data control flow. We demonstrate the ability of our methodology to understand various aspects of performance and performance bottlenecks on workflows and systems and motivate workflow optimizations. Nan Ding 0006, Brian Austin, Yang Liu 0179, Neil Mehta, Steven Farrell, Johannes P. Blaschke, Leonid Oliker, Hai Ah Nam, Nicholas J. Wright, Samuel Williams 0001 |
SC | 9 |
| 2024 | Understanding Data Movement Patterns in HPC: A NERSC Case StudyabstractScientific experiments are producing unprecedented volumes of data with real-time High Performance Computing (HPC) needs. Understanding and ensuring efficient data movement in these emerging data-intensive workloads is becoming critical for successful workflow execution. The need for end-to-end that integrates compute, network, and storage resources across facilities is resulting in a new integrated infrastructure paradigm. In this paper, we present an extensive analysis of three years of network traffic data from NERSC and identify critical data movement trends while detecting bottlenecks that significantly curtail transfer performance. Our results show that data movement patterns have shifted in the three years, and current infrastructure cannot sufficiently handle competing transfers, leading up to 30% throughput degradation for individual flows. In addition, we provide design recommendations for data movement management in future integrated research infrastructures that aim to reduce data transfer latency, reducing overall time to scientific results. Anna Giannakou, Damian Hazen, Bjoern Enders, Lavanya Ramakrishnan, Nicholas J. Wright |
SC | 5 |
| 2024 | Evaluating the potential of disaggregated memory systems for HPC applicationsabstractSummary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth. Nan Ding 0006, Pieter Maris, Hai Ah Nam, Taylor L. Groves, Muaaz Gul Awan, LeAnn Lindsey, Christopher S. Daley, Oguz Selvitopi, Leonid Oliker, Nicholas J. Wright, Samuel Williams 0001 |
Concurr. Comput. Pract. Exp. | 10 |
| 2024 | Architecture and performance of Perlmutter's 35 PB ClusterStor E1000 all-flash file systemabstractSummary NERSC's newest system, Perlmutter, features a 35 PB all‐flash Lustre file system built on HPE Cray ClusterStor E1000. We present its architecture, early performance figures, and performance considerations unique to this architecture. We demonstrate the performance of E1000 OSSes through low‐level Lustre tests that achieve over 90% of the theoretical bandwidth of the SSDs at the OST and LNet levels. We also show end‐to‐end performance for both traditional dimensions of I/O performance (peak bulk‐synchronous bandwidth) and nonoptimal workloads endemic to production computing (small, incoherent I/Os at random offsets) and compare them to NERSC's previous system, Cori, to illustrate that Perlmutter achieves the performance of a burst buffer and the resilience of a scratch file system. Finally, we discuss performance considerations unique to all‐flash Lustre and present ways in which users and HPC facilities can adjust their I/O patterns and operations to make optimal use of such architectures. Glenn K. Lockwood, Alberto Chiusole, Lisa Gerhardt, Kirill Lozinskiy, David Paul, Nicholas J. Wright |
Concurr. Comput. Pract. Exp. | 6 |
| 2023 | Performance-Aware Energy-Efficient GPU Frequency Selection using DNN-based ModelsabstractEnergy efficiency will be important in future accelerator-based HPC systems for sustainability and to improve overall performance. This study proposes a deep neural network (DNN)-based learning model for execution time and power consumption of workloads across GPUs DVFS design space. Micro-architectural data obtained by running SPEC-ACCEL, DGEMM, and STREAM benchmarks are used for model training. These features are consistent for a workload unaffected by frequency and input size reducing the data required significantly. For real-world applications - LAMMPS, NAMD, GROMACS, LSTM, BERT, and ResNet50 power and time models show 89% – 98% accuracy on NVIDIA Ampere. Multi-objective functions help select optimal frequencies that lower power and minimize performance impact showing maximum energy savings of 27% at a performance loss of 1.8%. The same models trained on Ampere showed an accuracy of greater than 93% on an NVIDIA Volta, thereby demonstrating model portability across architectures. Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001 |
ICPP | 4 |
| 2023 | Not all applications have boring communication patterns: Profiling message matching with BMMabstractSummary Message matching within MPI is an important performance consideration for applications that utilize two‐sided semantics. In this work, we present an instrumentation of the CrayMPI library that allows the collection of detailed message‐matching statistics as well as an implementation of hashed matching in software. We use this functionality to profile key DOE applications with complex communication patterns to determine under what circumstances an application might benefit from hardware offload capabilities within the NIC to accelerate message matching. We find that there are several applications and libraries that exhibit sufficiently long match list lengths to motivate a Binned Message Matching approach. Taylor L. Groves, Naveen Ravichandrasekaran, Brandon Cook 0001, Noel Keen, David Trebotich, Nicholas J. Wright, Robert Alverson, Duncan Roweth, Keith D. Underwood |
Concurr. Comput. Pract. Exp. | 6 |
| 2023 | An automated and portable method for selecting an optimal GPU frequency
Ghazanfar Ali, Mert Side, Sridutt Bhalachandra, Nicholas J. Wright, Yong Chen 0001 |
Future Gener. Comput. Syst. | 4 |
| 2022 | FPGA-based HPC accelerators: An evaluation on performance and energy efficiencyabstractAbstract Hardware specialization is a promising direction for the future of digital computing. Reconfigurable technologies enable hardware specialization with modest non‐recurring engineering cost, but their performance and energy efficiency compared to state‐of‐the‐art processor architectures remain an open question. In this article, we use FPGAs to evaluate the benefits of building specialized hardware for numerical kernels found in scientific applications. In order to properly evaluate performance, we not only compare Intel Arria 10 and Xilinx U280 performance against Intel Xeon, Intel Xeon Phi, and NVIDIA V100 GPUs, but we also extend the Empirical Roofline Toolkit (ERT) to FPGAs in order to assess our results in terms of the Roofline model. We show design optimization and tuning techniques for peak FPGA performance at reasonable hardware usage and power consumption. As FPGA peak performance is known to be far less than that of a GPU, we also benchmark the energy efficiency of each platform for the scientific kernels comparing against microbenchmark and technological limits. Results show that while FPGAs struggle to compete in absolute terms with GPUs on memory‐ and compute‐intensive kernels, they require far less power and can deliver nearly the same energy efficiency. Tan Nguyen 0001, Colin MacLean, Marco Siracusa, Douglas Doerfler, Nicholas J. Wright, Samuel Williams 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | Non-recurring engineering (NRE) best practices: a case study with the NERSC/NVIDIA OpenMP contractabstractThe NERSC supercomputer, Perlmutter, consists of AMD CPUs and NVIDIA GPUs. NERSC users expect to be able to use OpenMP to take advantage of the highly capable GPUs. This paper describes how NERSC/NVIDIA constructed a Non-Recurring Engineering (NRE) contract to add OpenMP GPU-offload support to the NVIDIA HPC compilers. The paper describes how the contract incorporated the strengths of both parties and encouraged collaboration to improve the quality of the final deliverable. We include our best practices and how this particular contract took into account emerging OpenMP specifications, NERSC workload requirements, and how to use OpenMP most efficiently on GPU hardware. This paper includes OpenMP application performance results obtained with the NVIDIA compilers distributed in the NVIDIA HPC SDK. Christopher S. Daley, Annemarie Southwell, Rahulkumar Gayatri, Scott Biersdorfff, Craig Toepfer, Güray Özen, Nicholas J. Wright |
SC | 7 |
| 2020 | Quantifying the impact of network congestion on application performance and network metricsabstractIn modern high-performance computing (HPC) systems, network congestion is an important factor that contributes to performance degradation. However, how network congestion impacts application performance is not fully understood. As Aries network, a recent HPC network architecture featuring a dragonfly topology, is equipped with network counters measuring packet transmission statistics on each router, these network metrics can potentially be utilized to understand network performance. In this work, by experiments on a large HPC system, we quantify the impact of network congestion on various applications' performance in terms of execution time, and we correlate application performance with network metrics. Our results demonstrate diverse impacts of network congestion: while applications with intensive MPI operations (such as HACC and MILC) suffer from more than 40% extension in their execution times under network congestion, applications with less intensive MPI operations (such as Graph500 and HPCG) are mostly not affected. We also demonstrate that a stall-to-flit ratio metric derived from Aries network counters is positively correlated with performance degradation and, thus, this metric can serve as an indicator of network congestion in HPC systems. Yijia Zhang 0002, Taylor L. Groves, Brandon Cook 0001, Nicholas J. Wright, Ayse K. Coskun |
CLUSTER | 4 |
| 2020 | Uncovering Access, Reuse, and Sharing Characteristics of I/O-Intensive Files on Large-Scale Production HPC Systems
Tirthak Patel, Surendra Byna, Glenn K. Lockwood, Nicholas J. Wright, Philip H. Carns, Robert B. Ross, Devesh Tiwari |
FAST | 4 |
| 2020 | Performance characterization of scientific workflows for the optimal use of Burst Buffers
Christopher S. Daley, Devarshi Ghoshal, Glenn K. Lockwood, Sudip S. Dosanjh, Lavanya Ramakrishnan, Nicholas J. Wright |
Future Gener. Comput. Syst. | 6 |
| 2019 | A Zoom-in Analysis of I/O Logs to Detect Root Causes of I/O Performance BottlenecksabstractScientific applications frequently spend a large fraction of their execution time in reading and writing data on parallel file systems. Identifying these I/O performance bottlenecks and attributing root causes are critical steps toward devising optimization strategies. Several existing studies analyze I/O logs of a set of benchmarks or applications that were run with controlled behaviors. However, there is still a lack of general approach that systematically identifies I/O performance bottlenecks for applications running "in the wild" on production systems. In this study, we have developed an analysis approach of "zooming in" from platform-wide to application-wide to job-level I/O logs for identifying I/O bottlenecks in arbitrary scientific applications. We analyze the logs collected on a Cray XC40 system in production over a two-month period. This study results in several insights for application developers to use in optimizing I/O behavior. Surendra Byna, Glenn K. Lockwood, Shane Snyder, Philip H. Carns, Sunggon Kim, Nicholas J. Wright |
CCGRID | 7 |
| 2019 | GPCNeT: designing a benchmark suite for inducing and measuring contention in HPC networksabstractNetwork congestion is one of the biggest problems facing HPC systems today, affecting system throughput, performance, user experience, and reproducibility. Congestion manifests as run-to-run variability due to contention for shared resources (e.g., filesystems) or routes between compute endpoints. Despite its significance, current network benchmarks fail to proxy the real-world network utilization seen on congested systems. We propose a new open-source benchmark suite called the Global Performance and Congestion Network Tests (GPCNeT) to advance the state of the practice in this area. The guiding principles used in designing GPCNeT are described and the methodology employed to maximize its utility is presented. The capabilities of GPCNeT are evaluated by analyzing results from several world's largest HPC systems, including an evaluation of congestion management on a next-generation network. The results show that systems of all technologies and scales are susceptible to congestion and this work motivates the need for congestion control in next-generation networks. Sudheer Chunduri, Taylor L. Groves, Peter Mendygral, Brian Austin, Jacob Balma, Krishna Kandalla, Kalyan Kumaran, Glenn K. Lockwood, Scott Parker, Steven Warren, Nathan Wichmann, Nicholas J. Wright |
SC | 12 |
| 2018 | IOMiner: Large-Scale Analytics Framework for Gaining Knowledge from I/O LogsabstractThe following topics are dealt with: parallel processing; message passing; parallel machines; application program interfaces; multiprocessing systems; resource allocation; scheduling; cache storage; data analysis; graphics processing units. Shane Snyder, Glenn K. Lockwood, Philip H. Carns, Nicholas J. Wright, Surendra Byna |
CLUSTER | 5 |
| 2018 | A year in the life of a parallel file system
Glenn K. Lockwood, Shane Snyder, Surendra Byna, Philip H. Carns, Nicholas J. Wright |
SC | 6 |
| 2017 | Understanding Performance Variability on the Aries Dragonfly NetworkabstractThis work evaluates performance variability in the Cray Aries dragonfly network and characterizes its impact on MPI Allreduce. The execution time of Allreduce is limited by the performance of the slowest participating process, which can vary by more than an order of magnitude. We utilize counters from the network routers to provide a better understanding of how competing workloads can influence performance. Specifically, we examine the relationships between message size, process counts, Aries counters and the Allreduce communication-time. Our results suggest that competing traffic from other jobs can significantly impact performance on the Aries Dragonfly Network. Furthermore, we show that Aries network counters are a valuable tool, explaining up to 70% of the performance variability for our experiments on a large-scale production system. Taylor L. Groves, Yizi Gu, Nicholas J. Wright |
CLUSTER | 3 |
| 2014 | Maximizing Throughput on a Dragonfly NetworkabstractInterconnection networks are a critical resource for large supercomputers. The dragonfly topology, which provides a low network diameter and large bisection bandwidth, is being explored as a promising option for building multi-Petaflop's and Exaflop's systems. Unlike the extensively studied torus networks, the best choices of message routing and job placement strategies for the dragonfly topology are not well understood. This paper aims at analyzing the behavior of a machine built using a dragonfly network for various routing strategies, job placement policies, and application communication patterns. Our study is based on a novel model that predicts traffic on individual links for direct, indirect, and adaptive routing strategies. We analyze results for individual communication patterns and some common parallel job workloads. The predictions presented in this paper are for a 100+ Petaflop's prototype machine with 92,160 high radix routers and 8.8 million cores. Abhinav Bhatele, Xiang Ni, Nicholas J. Wright, Laxmikant V. Kalé |
SC | 4 |
| 2010 | Performance Analysis of High Performance Computing Applications on the Amazon Web Services CloudabstractCloud computing has seen tremendous growth, particularly for commercial web applications. The on-demand, pay-as-you-go model creates a flexible and cost-effective means to access compute resources. For these reasons, the scientific computing community has shown increasing interest in exploring cloud computing. However, the underlying implementation and performance of clouds are very different from those at traditional supercomputing centers. It is therefore critical to evaluate the performance of HPC applications in today's cloud environments to understand the tradeoffs inherent in migrating to the cloud. This work represents the most comprehensive evaluation to date comparing conventional HPC platforms to Amazon EC2, using real applications representative of the workload at a typical supercomputing center. Overall results indicate that EC2 is six times slower than a typical mid-range Linux cluster, and twenty times slower than a modern HPC system. The interconnect on the EC2 cloud platform severely limits performance and causes significant variability. Keith R. Jackson, Lavanya Ramakrishnan, Krishna Muriki, Shane Canon, Shreyas Cholia, John Shalf, Harvey J. Wasserman, Nicholas J. Wright |
CloudCom | 8 |
| 2010 | Effective Performance Measurement at Petascale Using IPMabstractAs supercomputers are being built from an ever increasing number of processing elements, the effort required to achieve a substantial fraction of the system peak performance is continuously growing. Tools are needed that give developers and computing center staff holistic indicators about the resource consumption of applications and potential performance pitfalls at scale. To use the full potential of a supercomputer today, applications must incorporate multilevel parallelism (threading and message passing) and carefully orchestrate file I/O. As a consequence, performance tools must also be able to monitor these system components in an integrated way and at the full machine scales. We present IPM, a modularized monitoring approach for MPI, Open MP, file I/O, and other event sources. We describe its implementation design principles, which are targeted for efficiency and minimal application perturbation, and present an application study of using IPM at scale. Karl Fürlinger, Nicholas J. Wright, David Skinner |
ICPADS | 2 |
| 2010 | Parallel I/O performance: From events to ensemblesabstractParallel I/O is fast becoming a bottleneck to the research agendas of many users of extreme scale parallel computers. The principle cause of this is the concurrency explosion of high-end computation, coupled with the complexity of providing parallel file systems that perform reliably at such scales. More than just being a bottleneck, parallel I/O performance at scale is notoriously variable, being influenced by numerous factors inside and outside the application, thus making it extremely difficult to isolate cause and effect for performance events. In this paper, we propose a statistical approach to understanding I/O performance that moves from the analysis of performance events to the exploration of performance ensembles. Using this methodology, we examine two I/O-intensive scientific computations from cosmology and climate science, and demonstrate that our approach can identify application and middleware performance deficiencies - resulting in more than 4× run time improvement for both examined applications. Andrew Uselton, Mark Howison, Nicholas J. Wright, David Skinner, Noel Keen, John Shalf, Karen L. Karavanic, Leonid Oliker |
IPDPS | 3 |
| 2008 | Modeling and predicting application performance on parallel computers using HPC challenge benchmarksabstractA method is presented for modeling application performance on parallel computers in terms of the performance of microkernels from the HPC Challenge benchmarks. Specifically, the application run time is expressed as a linear combination of inverse speeds and latencies from microkernels or system characteristics. The model parameters are obtained by an automated series of least squares fits using backward elimination to ensure statistical significance. If necessary, outliers are deleted to ensure that the final fit is robust. Typically three or four terms appear in each model: at most one each for floating-point speed, memory bandwidth, interconnect bandwidth, and interconnect latency. Such models allow prediction of application performance on future computers from easier-to-make predictions of microkernel performance. The method was used to build models for four benchmark problems involving the PARATEC and MILC scientific applications. These models not only describe performance well on the ten computers used to build the models, but also do a good job of predicting performance on three additional computers with newer design features. For the four application benchmark problems with six predictions each, the relative root mean squared error in the predicted run times varies between 13 and 16%. The method was also used to build models for the HPL and G-FFTE benchmarks in HPCC, including functional dependences on problem size and core count from complexity analysis. The model for HPL predicts performance even better than the application models do, while the model for G-FFTE systematically underpredicts run times. Wayne Pfeiffer, Nicholas J. Wright |
IPDPS | 2 |
| 2007 | WRF nature runabstractThe Weather Research and Forecast (WRF) model is a limited-area model of the atmosphere for mesoscale research and operational numerical weather prediction (NWP). A petascale problem is a WRF nature run that provides very high-resolution "truth" against which more coarse simulations or perturbation runs may be compared for purposes of studying predictability, stochastic parameterization, and fundamental dynamics. We carried out a nature run involving an idealized high resolution rotating fluid on the hemisphere to investigate scales that span the k-3 to k-5/3 kinetic energy spectral transition of the observed atmosphere using 65,536 processors of the BG/L machine at LLNL. We worked through issues of parallel I/O and scalability. The primary result is not just the scalability and high Tflops number, but an important step towards understanding weather predictability at high resolution. John Michalakes, Josh Hacker, Richard Loft, Michael O. McCracken, Allan Snavely, Nicholas J. Wright, Thomas E. Spelce, Brent C. Gorda, Robert Walkup |
SC | 6 |