Scott Pakin

dblp:01/397 · DBLP profile ↗
← Back
49ranked-venue papers
10as first author
9since 2021 · last 2024
0000-0002-5220-1985ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 44 · 9 first-author · 7 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale System
abstract
Molecular dynamics (MD) simulations have transformed our understanding of the nanoscale, driving breakthroughs in materials science, computational chemistry, and several other fields, including biophysics and drug design. Even on exascale supercomputers, however, runtimes are excessive for systems and timescales of scientific interest. Here, we demonstrate strong scaling of MD simulations on the Cerebras Wafer-Scale Engine. By dedicating a processor core for each simulated atom, we demonstrate a 457-fold improvement in timesteps per second versus the Frontier GPU-based Exascale platform, along with a large improvement in timesteps per unit energy. Reducing every year of runtime to less than a day unlocks currently inaccessible timescales of slow microstructure transformation processes that are critical for understanding material behavior and function.Our dataflow algorithm runs Embedded Atom Method (EAM) simulations at rates over 699k timesteps per second for problems with up to 800k atoms. This demonstrated performance is unprecedented for general-purpose processing cores.
Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian, Ilya Sharapov, Aidan P. Thompson, Delyan Kalchev, Danny Perez, Robert Schreiber, Scott Pakin, Edgar A. León, James H. Laros III, Michael James 0002, Sivasankaran Rajamanickam
SC10
2024 Quantum optimization algorithms: Energetic implications
abstract
Summary Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. In this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least to more energy efficient than their classical counterparts; why feedback‐based quantum optimization is energy‐inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of 20. Finally, we use the NchooseK high‐level programming model to run optimization problems on both gate‐based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate‐based counterparts.
Rolando P. Hong Enriquez, Rosa M. Badia, Barbara M. Chapman, Kirk Bresniker, Scott Pakin, Alok Mishra 0002, Pedro Bruel, Aditya Dhakal, Gourav Rattihalli, Ninad Hogade, Eitan Frachtenberg, Dejan S. Milojicic
Concurr. Comput. Pract. Exp.5
2024 Quantum-centric supercomputing for materials science: A perspective on challenges and future directions
Yuri Alexeev, Maximilian Amsler, Marco Antonio Barroca, Sanzio Bassini, Torey Battelle, Daan Camps, David Casanova, Young Jay Choi, Fred Chong, Charles Chung, Christopher Codella, Antonio D. Córcoles, James Cruise, Alberto Di Meglio, Ivan Duran, Thomas Eckl, Sophia E. Economou, Stephan J. Eidenbenz, Bruce Elmegreen, Clyde Fare, Ismael Faro, Cristina Sanz Fernández, Rodrigo Neumann Barros Ferreira, Keisuke Fuji, Bryce Fuller, Laura Gagliardi, Giulia Galli, Jennifer R. Glick, Isacco Gobbi, Pranav Gokhale, Salvador de la Puente Gonzalez, Johannes Greiner, William Gropp, Michele Grossi, Emanuel Gull, Burns Healy, Matthew R. Hermes, Benchen Huang, Travis S. Humble, Nobuyasu Ito, Artur F. Izmaylov, Ali Javadi-Abhari, Douglas M. Jennewein, Shantenu Jha, Bert de Jong, Petar Jurcevic, William M. Kirby, Stefan Kister, Masahiro Kitagawa, Joel Klassen, Katherine Klymko, Kwangwon Koh, Masaaki Kondo, Doga Murat Kürkçüoglu, Krzysztof Kurowski, Teodoro Laino, Ryan Landfield, Matthew L. Leininger, Vicente Leyton-Ortega, Ang Li 0006, Meifeng Lin, Junyu Liu, Nicolás Lorente, André Luckow, Simon Martiel, Francisco Martín-Fernández, Margaret Martonosi, Claire Marvinney, Arcesio Castañeda Medina, Dirk Merten, Antonio Mezzacapo, Kristel Michielsen, Abhishek Mitra, Tushar Mittal, Kyungsun Moon, Joel Moore, Sarah Mostame, Mario Motta, Young-Hye Na, Yunseong Nam, Prineha Narang, Yu-ya Ohnishi, Daniele Ottaviani, Matthew Otten, Scott Pakin, Vincent R. Pascuzzi, Edwin Pednault, Tomasz Piontek, Jed W. Pitera, Patrick Rall, Gokul Subramanian Ravi, Niall Robertson, Matteo A. C. Rossi, Piotr Rydlichowski, Hoon Ryu, Georgy Samsonidze, Mitsuhisa Sato, Nishant Saurabh, Kunal Sharma, Soyoung Shin, George Slessman, Mathias Steiner, Iskandar Sitdikov, In-Saeng Suh, Eric D. Switzer, Joel Thompson, Synge Todo, Minh C. Tran, Dimitar Trenev, Christian Trott, Huan-Hsin Tseng, Norm M. Tubman, Esin Tureci, David García Valiñas, Sofia Vallecorsa, Christopher Wever, Konrad W. Wojciechowski, Xiaodi Wu 0001, Shinjae Yoo, Nobuyuki Yoshioka, Victor Wen-zhe Yu, Seiji Yunoki, Sergiy Zhuk, Dmitry Zubarev
Future Gener. Comput. Syst.87
2023 CLC: A cross-level program characterization method
abstract
Characterization of program execution plays a key role in performance improvement. There are numerous transformations applied to each step that a program takes on its lowering from source code to a compiler intermediate representation to machine language to microarchitecture-specific execution. The unpredictable benefit of each transformation step could lead a notionally superior algorithm to exhibit inferior performance once actually run, and it can be hard to discern which step in the transformation path contradicted the code developer’s assumptions. Conventional approaches to program-execution characterization consider the behavior after only a single one of those steps, which limits the information that can be provided to the user. To help address the issue of myopic views of program execution, this paper presents a novel cross-level characterization approach for understanding the behavior of program execution at different levels in the process of writing, compiling, and running a program. We show that this approach provides a richer view of the sources of performance gains and losses and helps identify program execution in a more accurate manner.
Li Tang 0007, Scott Pakin
Perform. Evaluation2
2022 Cross-Level Characterization of Program Behavior : (Extended Poster Abstract)
abstract
Program behavior can be defined as a collection of executions [1]. Program behavior strongly relates to actual program performance but can be complicated to be characterized and analyzed. Characterization is important as it helps better understand program behavior by measuring various operations a program performs. There are many existing techniques [2]–[7] for program characterization, which operate at different levels of instrumentation: source code, intermediate representation (IR), instruction set architecture (ISA), and CPU microarchitecture. Each of these levels provides different capabilities and limitations. In this paper, we introduce Cross-Level Characterization (CLC), an analysis of similarities and differences in resource counts as measured at each level of instrumentation during a program’s transformation from source code through execution on a specific microarchitecture.
Li Tang 0007, Scott Pakin
ISPASS2
2022 Cross-Level Characterization of Program Execution
abstract
Characterization of program execution plays a key role in performance improvement. There are numerous transformations applied to each step a program takes on its lowering from source code to a compiler intermediate representation to machine language to microarchitecture-specific execution. The unpredictable benefit of each transformation step could lead a notionally superior algorithm to exhibit inferior performance once actually run, and it can be opaque at what step in the transformation path contradicted the code developer's assumptions. However, conventional approaches to program execution characterization consider the behavior after only a single one of those steps, which limits the information that can be provided to the user. To help address the issue of myopic views of program execution, this paper presents a novel cross-level characterization approach for understanding the behavior of program execution at different levels in the process of writing, compiling, and running a program. We show that this approach provides a richer view of the sources of performance gains and losses and helps identify program execution in a more accurate manner.
Li Tang 0007, Scott Pakin
MASCOTS2
2022 Combining Hard and Soft Constraints in Quantum Constraint-Satisfaction Systems
abstract
This work presents a generalization of NchooseK, a constraint satisfaction system designed to target both quantum circuit devices and quantum annealing devices. Previously, NchooseK supported only hard constraints, which made it suitable for expressing problems in NP (e.g., 3-SAT) but not NP-hard problems (e.g., minimum vertex cover). In this paper we show how support for soft constraints can be added to the model and implementation, broadening the classes of problems that can be expressed elegantly in NchooseK without sacrificing portability across different quantum devices. Through a set of examples, we argue that this enhanced version of NchooseK enables problems to be expressed in a more concise, less error-prone manner than if these problems were encoded manually for quantum execution. We include an empirical evaluation of performance, scalability, and fidelity on both a large IBM Q system and a large D- Wave system.
Ellis Wilson, Frank Mueller 0001, Scott Pakin
SC3
2022 Guest Editorial: Special Section on Parallel and Distributed Computing Techniques for Non-Von Neumann Technologies
Scott Pakin, Christof Teuscher, Catherine D. Schuman
IEEE Trans. Parallel Distributed Syst.1
2022 Quantum Algorithm Implementations for Beginners
abstract
As quantum computers become available to the general public, the need has arisen to train a cohort of quantum programmers, many of whom have been developing classical computer programs for most of their careers. While currently available quantum computers have less than 100 qubits, quantum computing hardware is widely expected to grow in terms of qubit count, quality, and connectivity. This review aims at explaining the principles of quantum programming, which are quite different from classical programming, with straightforward algebra that makes understanding of the underlying fascinating quantum mechanical principles optional. We give an introduction to quantum computing algorithms and their implementation on real quantum hardware. We survey 20 different quantum algorithms, attempting to describe each in a succinct and self-contained fashion. We show how these algorithms can be implemented on IBM’s quantum computer, and in each case, we discuss the results of the implementation with respect to differences between the simulator and the actual hardware runs. This article introduces computer scientists, physicists, and engineers to quantum algorithms and provides a blueprint for their implementations.
Abhijith Jayakumar, Adetokunbo Adedoyin, John Ambrosiano, Petr M. Anisimov, William Casper, Gopinath Chennupati, Carleton Coffrin, Hristo N. Djidjev, David Gunter, Satish Karra, Nathan Lemons, Shizeng Lin, Alexander Malyzhenkov, David Mascarenas, Susan M. Mniszewski, Balasubramanya T. Nadiga, Daniel O'Malley, Diane Oyen, Scott Pakin, Lakshman Prasad, Randy Roberts, Phillip Romero, Nandakishore Santhi, Nikolai Sinitsyn, Pieter J. Swart, Jim Wendelberger, Boram Yoon, Richard J. Zamora, Wei Zhu 0011, Stephan J. Eidenbenz, Andreas Bärtschi, Patrick J. Coles, Marc Vuffray, Andrey Y. Lokhov
ACM Trans. Quantum Comput.19
2019 Targeting Classical Code to a Quantum Annealer
abstract
From a compiler's perspective, a quantum annealer represents a fundamentally different hardware target from a CPU, GPU, or other von Neumann architecture. Quantum annealers are special-purpose computers that use quantum effects to heuristically determine the set of Boolean variables that minimize a quadratic pseudo-Boolean function (an NP-hard problem). Natively programming such systems involves supplying them with a vector of function coefficients and receiving a vector of function-minimizing Booleans in return. The contribution of this work is to demonstrate how to compile conventional code into a minimization problem for solution on a quantum annealer. The resulting code can run either forward (from inputs to outputs) or backward (from outputs to inputs). We show how this capability can be exploited to simplify the expression and solution of problems in the NP complexity class.
Scott Pakin
ASPLOS1
2019 Implementing NChooseK on IBM Q Quantum Computer Systems
Harsh Khetawat, Ashlesha Atrey, George Li 0001, Frank Mueller 0001, Scott Pakin
RC5
2018 A Comparative Study of Topology Design Approaches for HPC Interconnects
abstract
The recent interconnect topology designs for High Performance Computing (HPC) systems have followed two directions, one characterized by low diameter and the other by high path diversity. The low diameter design focuses on building large networks with small diameters, guaranteeing one short path between each pair of nodes. Examples include Slim Fly and Dragonfly. The high path diversity design takes into account not only other topological metrics such as diameter but also path diversity between pairs of nodes. Examples include fat-tree, Random Regular Graph (RRG) and Generalized De Bruin Graph (GDBG). Topologies designed from these two approaches have distinct features and require very different routing schemes to exploit the network capacity. In this work, we study the performance-related topological features of representative topologies of the two design approaches, including Slim Fly, Dragonfly, RRG, and GDBG, and compare HPC application performance on these topologies with a set of routing schemes. The study uncovers new knowledge about the topologies designed by these two approaches. Findings of the study include (1) the load balance routing technique designed for low diameter topologies, known as the Universal Globally Adaptive Load-balanced routing (UGAL), can be effectively adapted for the high path diversity topologies, and (2) high path diversity topologies in general achieve higher performance than low diameter topologies for networks built by a similar number of the same type of switches.
Md Atiqul Mollah, Peyman Faizian, Md. Shafayat Rahman, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
CCGrid5
2018 Performance and Accuracy Trade-offs of HPC Application Modeling and Simulation
abstract
High Performance Computing (HPC) applications and systems are often studied through modeling and simulation at various granularities. As the size of HPC systems and applications and the cost of high fidelity simulation continue to grow, a good understanding of the trade-offs of the complexity and accuracy of HPC application modeling and simulation schemes can help balance the competing goals of accuracy and time. In this work, we investigate the complexity and accuracy trade-off using an MPI application modeling tool and an MPI application simulation tool. The performance and accuracy results of modeling and simulation of a large spectrum of HPC applications on three supercomputers are measured and compared. The results show that although modeling is often one to two orders of magnitude faster than simulation, it achieves within 5% of predicted application time in comparison to simulation for 85% of cases in our data set. We further enhance the modeling tool with a statistical model to predict whether simulation can yield significantly different results than modeling. The enhanced tool achieves a very high successful prediction rate of 93.2% on our dataset and is thus effective in determining whether modeling or simulation should be used.
Zhou Tong, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
IPDPS3
2018 Fast classification of MPI applications using Lamport's logical clocks
Zhou Tong, Scott Pakin, Michael Lang 0003, Xin Yuan 0001
J. Parallel Distributed Comput.2
2018 Random Regular Graph and Generalized De Bruijn Graph with k-Shortest Path Routing
abstract
The Random regular graph (RRG) has recently been proposed as an interconnect topology for future large scale data centers and HPC clusters. An RRG is a special case of directed regular graph (DRG) where each link is unidirectional and all nodes have the same number of incoming and outgoing links. In this work, we establish bounds for DRGs on diameter, average k-shortest path length, and a load balancing property with k-shortest path routing, and use these bounds to evaluate RRGs. The results indicate that an RRG with k-shortest path routing is not ideal in terms of diameter and load balancing. We further consider the Generalized De Bruijn Graph (GDBG), a deterministic DRG, and prove that for most network configurations, a GDBG is near optimal in terms of diameter, average k-shortest path length, and load balancing with a k-shortest path routing scheme. Finally, we use modeling and simulation to exploit the strengths and weaknesses of RRGs for different traffic conditions by comparing RRGs with GDBGs.
Peyman Faizian, Md Atiqul Mollah, Xin Yuan 0001, Zaid Salamah A. Alzaid, Scott Pakin, Michael Lang 0003
IEEE Trans. Parallel Distributed Syst.5
2018 Rapid Calculation of Max-Min Fair Rates for Multi-Commodity Flows in Fat-Tree Networks
abstract
Max-min fairness is often used in the performance modeling of interconnection networks. Existing methods to compute max-min fair rates for multi-commodity flows have high complexity and are computationally infeasible for large networks. In this work, we show that by considering topological features, this problem can be solved efficiently for the fat-tree topology that is widely used in data centers and high performance compute clusters. Several efficient new algorithms are developed for this problem, including a parallel algorithm that can take advantage of multi-core and shared-memory architectures. Using these algorithms, we demonstrate that it is possible to find the max-min fair rate allocation for multi-commodity flows in fat-tree networks that support tens of thousands of nodes. We evaluate the run-time performance of the proposed algorithms and show improvement in orders of magnitude over the previously best known method. We further demonstrate a new application of max-min fair rate allocation that is only computationally feasible using our new algorithms.
Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
IEEE Trans. Parallel Distributed Syst.3
2018 Performing fully parallel constraint logic programming on a quantum annealer
abstract
Abstract Aquantum annealerexploits quantum effects to solve a particular type of optimization problem. The advantage of this specialized hardware is that it effectively considers all possible solutions in parallel, thereby potentially outperforming classical computing systems. However, despite quantum annealers having recently become commercially available, there are relatively few high-level programming models that target these devices. In this article, we show how to compile a subset of Prolog enhanced with support for constraint logic programming into a two-local Ising-model Hamiltonian suitable for execution on a quantum annealer. In particular, we describe the series of transformations one can apply to convert constraint logic programs expressed in Prolog into an executable form that bears virtually no resemblance to a classical machine model yet that evaluates the specified constraints in a fully parallel manner. We evaluate our efforts on a 1,095-qubit D-Wave 2X quantum annealer and describe the approach's associated capabilities and shortcomings.
Scott Pakin
Theory Pract. Log. Program.1
2017 Characterizing and Modeling Power and Energy for Extreme-Scale In-Situ Visualization
abstract
Plans for exascale computing have identified power and energy as looming problems for simulations running at that scale. In particular, writing to disk all the data generated by these simulations is becoming prohibitively expensive due to the energy consumption of the supercomputer while it idles waiting for data to be written to permanent storage. In addition, the power cost of data movement is also steadily increasing. A solution to this problem is to write only a small fraction of the data generated while still maintaining the cognitive fidelity of the visualization. With domain scientists increasingly amenable towards adopting an in-situ framework that can identify and extract valuable data from extremely large simulation results and write them to permanent storage as compact images, a large-scale simulation will commit to disk a reduced dataset of data extracts that will be much smaller than the raw results, resulting in a savings in both power and energy. The goal of this paper is two-fold: (i) to understand the role of in-situ techniques in combating power and energy issues of extreme-scale visualization and (ii) to create a model for performance, power, energy, and storage to facilitate what-if analysis. Our experiments on a specially instrumented, dedicated 150-node cluster show that while it is difficult to achieve power savings in practice using in-situ techniques, applications can achieve significant energy savings due to shorter write times for in-situ visualization. We present a characterization of power and energy for in-situ visualization; an application-aware, architecture-specific methodology for modeling and analysis of such in-situ workflows; and results that uncover indirect power savings in visualization workflows for high-performance computing (HPC).
Vignesh Adhinarayanan, Wu-chun Feng, David H. Rogers 0001, James P. Ahrens, Scott Pakin
IPDPS5
2016 Random Regular Graph and Generalized De Bruijn Graph with k-Shortest Path Routing
abstract
Random regular graph (RRG) has recently been proposed as an interconnect topology for future large scale data centers and HPC clusters. While various studies have been performed, this topology is still not well understood. RRG is a special case of directed regular graph (DRG) where each link is unidirectional and all nodes have the same number of incoming and outgoing links. In this work, we establish bounds for DRG on diameter, average k-shortest path length, and a load balancing property with k-shortest path routing, and use these bounds to evaluate RRG. The results indicate that RRG with k-shortest path routing is not ideal in terms of diameter and load balancing. We further consider the Generalized De Bruijn Graph (GDBG), a deterministic DRG, and prove that for most network configurations, GDBG is near optimal in terms of diameter, average k-shortest path length, and load balancing with a k-shortest path routing scheme. Finally, we explore the strengths and weaknesses of RRG for different traffic conditions by comparing RRG with GDBG.
Peyman Faizian, Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
IPDPS4
2016 Fast Classification of MPI Applications Using Lamport's Logical Clocks
abstract
We present a novel trace-based analysis tool that rapidly classifies an MPI application as bandwidth-bound, latency-bound, load-imbalance-bound, or computation-bound for different interconnection networks. The tool uses an extension of Lamport's logical clock to track application progress in the trace replay. Ithas two unique features. First, it predicts application performance for many latency and bandwidth parameters from a single replay of the trace. Second, it infers the performance characteristics of an application and classifies the application using the predicted performance trend for a range of network configurations instead of using the predicted performance for a particular network configuration. We describe the techniques used in the tool and its design and implementation, and report our performance study of the tool and our experience with classifying nine applications and mini-apps from the DOE Design Forward project as well as the NAS Parallel Benchmarks.
Zhou Tong, Scott Pakin, Michael Lang 0003, Xin Yuan 0001
IPDPS2
2016 Power usage of production supercomputers and production workloads
abstract
Summary Power is becoming an increasingly important concern for large supercomputer centers. However, to date, there have been a dearth of studies of power usage ‘in the wild’—on production supercomputers running production workloads. In this paper, we present the initial results of a project to characterize the power usage of the three Top500 supercomputers at Los Alamos National Laboratory: Cielo, Roadrunner, and Luna (#15, #19, and #47, respectively, on the June 2012 Top500 list). Power measurements taken both at the switchboard level and within the compute racks are presented and discussed. Some noteworthy results of this study are that (1) variability in power consumption differs across architectures, even when running a similar workload and (2) Los Alamos National Laboratory's scientific workload draws, on average, only 70–75% of LINPACK power and only 40–55% of nameplate power, implying that power capping may enable a substantial reduction in power and cooling infrastructure while impacting comparatively few applications. Copyright © 2013 John Wiley & Sons, Ltd.
Scott Pakin, Curtis B. Storlie, Michael Lang 0003, Bob Fields, Eloy E. Romero Jr., Craig Idler, Sarah Ellen Michalak, Hugh Greenberg, Josip Loncaric, Randal Rheinheimer, Gary Grider, Joanne Wendelberger
Concurr. Comput. Pract. Exp.1
2016 TracSim: Simulating and scheduling trapped power capacity to maximize machine room throughput
Michael Lang 0003, Scott Pakin, Song Fu
Parallel Comput.3
2015 Fast Calculation of Max-Min Fair Rates for Multi-commodity Flows in Fat-Tree Networks
abstract
Max-min fairness is often used in the performance modeling of interconnection networks. Existing methods to compute max-min fair rates for multi-commodity flows have high complexity and are computationally infeasible for large networks. In this work, we show that by considering topological features, this problem can be solved efficiently for the fat-tree topology that is widely used in data centers and high performance computing clusters. Using two new algorithms that we developed, we demonstrate it is possible to find the max-min fair rate allocation for multi-commodity flows in fat-tree networks that support tens of thousands of nodes. We evaluate the run-time performance of the proposed algorithms and demonstrate an application.
Md Atiqul Mollah, Xin Yuan 0001, Scott Pakin, Michael Lang 0003
CLUSTER3
2014 Computational Co-design of a Multiscale Plasma Application: A Process and Initial Results
abstract
As computer architectures become increasingly heterogeneous the need for algorithms and applications that can exploit these new architectures grows more pressing. This paper demonstrates that co-designing a multi-architecture, multi-scale, highly optimized framework with its associated plasma-physics application can provide both portability across CPUs and accelerators and high performance. Our framework utilizes multiple abstraction layers in order to maximize code reuse between architectures while providing low-level abstractions to incorporate architecture-specific optimizations such as vectorization or hardware fused multiply-add. We describe a co-design process used to enable a plasma physics application to scale well to large systems while also improving on both the accuracy and speed of the simulations. Optimized multi-core results will be presented to demonstrate ability to isolate large amounts of computational work with minimal communication.
Joshua Payne, Dana A. Knoll, Allen McPherson, William T. Taitano, Luis Chacón, Guangye Chen, Scott Pakin
IPDPS7
2014 LFTI: A New Performance Metric for Assessing Interconnect Designs for Extreme-Scale HPC Systems
abstract
Traditionally, interconnect performance is either characterized by simple topological parameters such as bisection bandwidth or studied through simulation that gives detailed performance information for the scenarios simulated. Neither of these approaches provides a good performance overview for extreme-scale interconnects. The topological parameters are not directly related to application level communication performance while the simulation complexity limits the number of scenarios that can be investigated. In this work, we propose a new performance metric, called LANL-FSU Throughput Indices (LFTI), for characterizing the throughput performance of interconnect designs. LFTI combines the simplicity of topological parameters and the accuracy of simulation: like topological parameters, LFTI can be derived from interconnect specification, at the same time, it directly reflects the application level communication performance. Moreover, in cases when the theoretical throughput for each communication pattern can be modeled efficiently for an interconnect, LFTI for the interconnect can be computed efficiently. These features potentially allow LFTI to be used for rapid and comprehensive evaluation and comparison of extreme-scale interconnect designs. We demonstrate the effectiveness of LFTI by using it to evaluate and explore the design space of a number of large-scale interconnect designs.
Xin Yuan 0001, Santosh Mahapatra, Michael Lang 0003, Scott Pakin
IPDPS4
2014 Static load-balanced routing for slimmed fat-trees
Xin Yuan 0001, Santosh Mahapatra, Michael Lang 0003, Scott Pakin
J. Parallel Distributed Comput.4
2013 Topic 2: Performance Prediction and Evaluation - (Introduction)
Adolfy Hoisie, Michael Gerndt, Shajulin Benedict, Thomas Fahringer, Vladimir Getov, Scott Pakin
Euro-Par6
2013 Exploring power behaviors and trade-offs of in-situ data analytics
abstract
As scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems.
Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky
SC11
2013 A new routing scheme for Jellyfish and its performance with HPC workloads
abstract
The jellyfish topology where switches are connected using a random graph has recently been proposed for large scale data-center networks. It has been shown to offer higher bisection bandwidth and better permutation throughput than the corresponding fat-tree topology with a similar cost. In this work, we propose a new routing scheme for jellyfish that out-performs existing schemes by more effectively exploiting the path diversity, and comprehensively compare the performance of jellyfish and fat-tree topologies with HPC workloads. The results indicate that both jellyfish and fat-tree topologies offer comparable high performance for HPC workloads on systems that can be realized by 3-level fat-trees using the current technology and the corresponding jellyfish topologies with similar costs. Fat-trees are more effective for smaller systems while jellyfish is more scalable.
Xin Yuan 0001, Santosh Mahapatra, Wickus Nienaber, Scott Pakin, Michael Lang 0003
SC4
2012 Special issue on Communication Architectures for Scalable Systems
José Flich, Scott Pakin, Craig B. Stunkel
J. Parallel Distributed Comput.2
2011 Automatic generation of executable communication specifications from parallel applications
abstract
Portable parallel benchmarks are widely used and highly effective for (a) the evaluation, analysis and procurement of high-performance computing (HPC) systems and (b) quantifying the potential benefits of porting applications for new hardware platforms. Yet, past techniques to synthetically parametrized hand-coded HPC benchmarks prove insufficient for today's rapidly-evolving scientific codes particularly when subject to multi-scale science modeling or when utilizing domain-specific libraries.
Xing Wu 0004, Frank Mueller 0001, Scott Pakin
ICS3
2011 Adapting wave-front algorithms to efficiently utilize systems with deep communication hierarchies
Darren J. Kerbyson, Michael Lang 0003, Scott Pakin
Parallel Comput.3
2009 Application profiling on Cell-based clusters
abstract
In this paper, we present a methodology for profiling parallel applications executing on the IBM PowerXCell 8i (commonly referred to as the ldquoCellrdquo processor). Specifically, we examine Cell-centric MPI programs on hybrid clusters containing multiple Opteron and Cell processors per node such as those used in the petascale Roadrunner system. Our implementation incurs less than 3.2 mus of overhead per profile call while efficiently utilizing the limited local store of the Cell's SPE cores. We demonstrate the use of our profiler on a cluster of hybrid nodes running a suite of scientific applications. Our analyses of inter-SPE communication (across the entire cluster) and function call patterns provide valuable information that can be used to optimize application performance.
Hikmet Dursun, Kevin J. Barker, Darren J. Kerbyson, Scott Pakin
IPDPS4
2008 Experiences in scaling scientific applications on current-generation quad-core processors
abstract
In this work we present an initial performance evaluation of AMD and Intel's first quad-core processor offerings: the AMD Barcelona and the Intel Xeon X7350. We examine the suitability of these processors in quad-socket compute nodes as building blocks for large-scale scientific computing clusters. Our analysis of intra-processor and intra-node scalability of microbenchmarks and a range of large- scale scientific applications indicates that quad-core processors can deliver an improvement in performance of up to 4x per processor but is heavily dependent on the workload being processed. While the Intel processor has a higher clock rate and peak performance, the AMD processor has higher memory bandwidth and intra-node scalability. The scientific applications we analyzed exhibit a range of performance improvements from only 3x up to the full 16x speed-up over a single core. Also, we note that the maximum node performance is not necessarily achieved by using all 16 cores.
Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho
IPDPS6
2008 Receiver-initiated message passing over RDMA Networks
abstract
Providing point-to-point messaging-passing semantics atop Put/Get hardware traditionally involves implementing a protocol comprising three network latencies. In this paper, we analyze the performance of an alternative implementation approach - receiver-initiated message passing - that eliminates one of the three network latencies. Performance measurements taken on the Cell Broadband Engine indicate that receiver-initiated message passing exhibits substantially lower latency than standard, sender-initiated message passing.
Scott Pakin
IPDPS1
2008 Entering the petaflop era: the architecture and performance of Roadrunner
abstract
Roadrunner is a 1.38 Pflop/s-peak (double precision) hybrid-architecture supercomputer developed by LANL and IBM. It contains 12,240 IBM PowerXCell 8i processors and 12,240 AMD Opteron cores in 3,060 compute nodes. Roadrunner is the first supercomputer to run Linpack at a sustained speed in excess of 1 Pflop/s. In this paper we present a detailed architectural description of Roadrunner and a detailed performance analysis of the system. A case study of optimizing the MPI-based application Sweep3D to exploit Roadrunner's hybrid architecture is also included. The performance of Sweep3D is compared to that of the code on a previous implementation of the Cell Broadband Engine architecture-the Cell BE-and on multi-core processors. Using validated performance models combined with Roadrunner-specific microbenchmarks we identify performance issues in the early pre-delivery system and infer how well the final Roadrunner configuration will perform once the system software stack has matured.
Kevin J. Barker, Kei Davis, Adolfy Hoisie, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, José Carlos Sancho
SC6
2007 Performance analysis of a user-level memory server
abstract
Large-scale parallel applications often produce immense quantities of data that need to be analyzed. To avoid performing repeated, costly disk accesses, analysis of large data sets generally requires a commensurately large amount of memory. While some data-analysis tools can easily be parallelized to distribute memory across a cluster, other tools are either difficult to parallelize or, in the case of simple data-analysis scripts with short lifespans, not worth the effort to parallelize. In this work, we present and analyze the performance of JumboMem, a simple, entirely user-level parallel program that enables unmodified sequential applications to access all of the memory in a cluster. Although there are many implementations of memory servers, all require either administrative privileges or program modifications. More importantly, no existing memory server has been evaluated on modern workstation clusters with high-speed networks, many nodes, and significant quantities of memory. This paper represents the first study of memory-server performance at supercomputing scales.
Scott Pakin, Greg Johnson
CLUSTER1
2007 The Design and Implementation of a Domain-Specific Language for Network Performance Testing
abstract
CONCEPTUAL is a toolset designed specifically to help measure the performance of high-speed interconnection networks such as those used in workstation clusters and parallel computers. It centers around a high-level domain-specific language, which makes it easy for a programmer to express, measure, and report the performance of complex communication patterns. The primary challenge in implementing a compiler for such a language is that the generated code must be extremely efficient so as not to misattribute overhead costs to the messaging library. At the same time, the language itself must not sacrifice expressiveness for compiler efficiency, or there would be little point in using a high-level language for performance testing. This paper describes the CONCEPTUAL language and the CONCEPTUAL compiler's novel code-generation framework. The language provides primitives for a wide variety of idioms needed for performance testing and emphasizes a readable syntax. The core code-generation technique, based on unrolling CONCEPTUAL programs into sequences of communication events, is simple yet enables the efficient implementation of a variety of high-level constructs. The paper further explains how CONCEPTUAL implements time-bounded loops - even those that comprise blocking communication - in the absence of a time-out mechanism as this is a somewhat unique language/implementation feature.
Scott Pakin
IEEE Trans. Parallel Distributed Syst.1
2006 A Performance Model of the Krak Hydrodynamics Application
abstract
We present an analytic performance model of a large-scale hydrodynamics code developed at Los Alamos National Laboratory. This modeling work is part of an ongoing effort to develop models and modeling techniques for large-scale codes and systems of interest to Los Alamos and the national laboratory community (Kerbyson et al., 2001). Krak (Burton, 1994) comprises over 270,000 lines of source code and is capable of executing on a large number of parallel processors. Developing an accurate model is complicated by the irregular partitioning of input spatial grid cells to processors and the various material properties assigned to each cell. Model development proceeds by separating inter-processor communication from computation and modeling each individually. In addition, several approximations concerning subgrid size, shape, and material composition are made which reduce modeling complexity without adversely impacting prediction accuracy. We validate our model on several spatial grid sizes and processor configurations and demonstrate an accuracy at the largest scale on 512 processors to within a 3% error
Kevin J. Barker, Scott Pakin, Darren J. Kerbyson
ICPP2
2006 Architecture - A performance comparison through benchmarking and modeling of three leading supercomputers: blue Gene/L, Red Storm, and Purple
abstract
This work provides a performance analysis of three leading supercomputers that have recently been deployed: Purple, Red Storm and Blue Gene/L. Each of these machines are architecturally diverse, with very different performance characteristics. Each contains over 10,000 processors and has a system peak of over 40 Teraflops. We analyze each system using a range of micro-benchmarks which include communication performance as well as quantifying the impact of the operating system. The achievable application performance is compared across the systems. The application performance is confirmed via the use of detailed application models which use the underlying performance characteristics as measured by the micro-benchmarks. We also compare the machines in a realistic production scenario in which each machine is used so as to maximize its memory usage with the applications executed in a weak-scaling mode. The results also help illustrate that achievable performance is not directly related to the peak performance.
Adolfy Hoisie, Greg Johnson, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin
SC5
2006 STORM: Scalable Resource Management for Large-Scale Parallel Computers
abstract
Although clusters are a popular form of high-performance computing, they remain more difficult to manage than sequential systems - or even symmetric multiprocessors. In this paper, we identify a small set of primitive mechanisms that are sufficiently general to be used as building blocks to solve a variety of resource-management problems. We then present STORM, a resource-management environment that embodies these mechanisms in a scalable, low-overhead, and efficient implementation. The key innovation behind STORM is a modular software architecture that reduces all resource management functionality to a small number of highly scalable mechanisms. These mechanisms simplify the integration of resource management with low-level network features. As a result of this design, STORM can launch large, parallel applications an order of magnitude faster than the best time reported in the literature and can gang-schedule a parallel application as fast as the node OS can schedule a sequential application. This paper describes the mechanisms and algorithms behind STORM and presents a detailed performance model that shows that STORM's performance can scale to thousands of nodes
Eitan Frachtenberg, Fabrizio Petrini, Juan Fernández Peinador, Scott Pakin
IEEE Trans. Computers4
2004 Reproducible Network Benchmarks with coNCePTuaL
Scott Pakin
Euro-Par1
2004 coNCePTuaL: A Network Correctness and Performance Testing Languag
abstract
Summary form only given. We introduce a new, domain-specific specification language called CONCEPTUAL. CONCEPTUAL enables the expression of sophisticated communication benchmarks and network validation tests in comparatively few lines of code. Besides helping programmers save time writing and debugging code, CONCEPTUAL addresses the important-but largely unrecognized-problem of benchmark opacity. Benchmark opacity refers to the current impracticality of presenting performance measurements in a manner that promotes reproducibility and independent evaluation of the results. For example, stating that a performance graph was produced by a "bandwidth" test says nothing about whether that test measures the data rate during a round-trip transmission or the average data rate over a number of back-to-back unidirectional messages; whether the benchmark preregisters buffers, sends warm-up messages, and/or preposts asynchronous receives before starting the clock; how many runs were performed and whether these were aggregated by taking the mean, median, or maximum; or, even whether a data unit such as "MB/s" indicates 10/sup 6/ or 2/sup 20/ bytes per second. Because CONCEPTUAL programs are terse, a benchmark's complete source code can be listed alongside performance results, making explicit all of the design decisions that went into the benchmark program. Because CONCEPTUAL's grammar is English-like, CONCEPTUAL programs can easily be understood by nonexperts. And because CONCEPTUAL is a high-level language, it can target a variety of messaging layers and networks, enabling fair and accurate performance comparisons.
Scott Pakin
IPDPS1
2004 A Performance and Scalability Analysis of the BlueGene/L Architecture
abstract
Based on a set of measurements done on the 512-node 500MHz prototype and early results on a 2048 node 700MHz BlueGene/L machine at IBM Watson, we present a performance and scalability analysis of the architecture from low-level characteristics to large-scale applications. In addition, we present predictions using our models for the performance of two representative applications from the ASC² workload on the full BlueGene/L configuration of 64K nodes. We have compared the measured values for several of the benchmarks in our suite against the predicted numbers from our performance models. In general, the error bars were relatively low. A comparison between the performance of BlueGene/L and the ASCI Q, the largest supercomputer in the US, is presented, also based on our predictive performance models.
Kei Davis, Adolfy Hoisie, Greg Johnson, Darren J. Kerbyson, Michael Lang 0003, Scott Pakin, Fabrizio Petrini
SC6
2003 The Case of the Missing Supercomputer Performance: Achieving Optimal Performance on the 8, 192 Processors of ASCI Q
abstract
In this paper we describe how we improved the effective performance of ASCI Q, the world's second-fastest supercomputer, to meet our expectations. Using an arsenal of performance-analysis techniques including analytical models, custom microbenchmarks, full applications, and simulators, we succeeded in observing a serious-but previously undetectable-performance problem. We identified the source of the problem, eliminated the problem, and 'closed the loop' by demonstrating improved application performance. We present our methodology and provide insight into performance analysis that is immediately applicable to other large-scale cluster-based supercomputers.
Fabrizio Petrini, Darren J. Kerbyson, Scott Pakin
SC3
2002 STORM: lightning-fast resource management
abstract
Although workstation clusters are a common platform for high-performance computing (HPC), they remain more difficult to manage than sequential systems or even symmetric multiprocessors. Furthermore, as cluster sizes increase, the quality of the resource-management subsystem — essentially, all of the code that runs on a cluster other than the applications — increasingly impacts application efficiency. In this paper, we present STORM, a resource-management framework designed for scalability and performance. The key innovation behind STORM is a software architecture that enables resource management to exploit low-level network features. As a result of this HPC-application-like design, STORM is orders of magnitude faster than the best reported results in the literature on two sample resource-management functions: job launching and process scheduling.
Eitan Frachtenberg, Fabrizio Petrini, Juan Fernández Peinador, Scott Pakin, Salvador Coll
SC4
1998 Efficient Layering for High Speed Communication: Fast Messages 2.x
abstract
The authors describe their experience designing, implementing, and evaluating two generations of the high performance communication library, Fast Messages (FM) for Myrinet. In FM 1.x, they designed a simple interface and provided guarantees of reliable and in-order delivery, and flow control. While this was a significant improvement over previous systems, it was not enough. Layering MPI atop FM 1.x showed that only about 20% of the FM 1.x bandwidth could be delivered to higher level communication APIs. The second generation communication layer, FM 2.0, addresses the identified problems, providing gather-scatter, interlayer scheduling, receiver flow control, as well as some convenient API features which simplify programming. FM 2.x can deliver 70-90% to higher level APIs such as MPI. This is especially impressive as the absolute bandwidths delivered have increased nearly fourfold to 70 MB/s. They describe general issues encountered in matching two communication layers, and the solutions as embodied in FM 2.x.
Mario Lauria, Scott Pakin, Andrew A. Chien
HPDC2
1998 Dynamic Coscheduling on Workstation Clusters
Patrick Sobalvarro, Scott Pakin, William E. Weihl, Andrew A. Chien
JSSPP2
1995 High Performance Messaging on Workstations: Illinois Fast Messages (FM) for Myrinet
Scott Pakin, Mario Lauria, Andrew A. Chien
SC1