EDBT 2026 Demo / reviewers in the wild / expert
Patrick H. Worley
dblp:67/4971
· DBLP profile ↗
21ranked-venue papers
7as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 7 first-authorSoftware engineering, systems software and programming languages · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
9 papers |
High-performance computing · 53% Performance modeling and evaluation · 36% Processor architecture and microarchitecture · 4% | |
| Software engineering, system software, and programming languages
1 paper |
Program verification · 100% |
Topics — the 14 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
benchmarking |
0.3 | 6 | 2011 | Early evaluation of IBM BlueGene/P · SC 2008 Cray XT4: an early evaluation for petascale scientific simulation · SC 2007 Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005 |
High-performance computing
performance optimization at scale |
0.2 | 3 | 2011 | Performance of the community earth system model · SC 2011 Scaling the unscalable: a case study on the AlphaServer SC · SC 2002 Early Evaluation of the Cray X1 · SC 2003 |
High-performance computing › scientific computing systems
climate modeling |
0.1 | 2 | 2011 | Performance of the community earth system model · SC 2011 Performance Tuning and Evaluation of a Parallel Community Climate Model · SC 1999 |
High-performance computing
scientific computing systems |
0.1 | 2 | 2011 | Performance of the community earth system model · SC 2011 Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005 |
High-performance computing › supercomputing
supercomputing systems |
0.1 | 2 | 2008 | Early evaluation of IBM BlueGene/P · SC 2008 Early evaluation of the IBM p690 · SC 2002 |
Performance modeling and evaluation › parallel performance evaluation
scientific application performance |
0.1 | 1 | 2005 | Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005 |
Processor architecture and microarchitecture
vector processor |
0.1 | 1 | 2005 | Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005 |
Performance modeling and evaluation › system-level analysis
architecture evaluation |
0.0 | 2 | 2003 | Early evaluation of the IBM p690 · SC 2002 Early Evaluation of the Cray X1 · SC 2003 |
High-performance computing
supercomputing |
0.0 | 1 | 2003 | Early Evaluation of the Cray X1 · SC 2003 |
Program verification › dynamic verification
runtime assertion checking |
0.0 | 1 | 2002 | Asserting performance expectations · SC 2002 |
Performance modeling and evaluation
performance analysis tools |
0.0 | 1 | 2002 | Scaling the unscalable: a case study on the AlphaServer SC · SC 2002 |
Embedded and real-time systems › runtime monitoring
runtime verification |
0.0 | 1 | 2002 | Asserting performance expectations · SC 2002 |
Performance modeling and evaluation
performance tuning |
0.0 | 1 | 1999 | Performance Tuning and Evaluation of a Parallel Community Climate Model · SC 1999 |
High-performance computing › scientific computing systems
atmospheric modeling |
0.0 | 1 | 2005 | Leading Computational Methods on Scalar and Vector HEC Platforms · SC 2005 |
Methods — techniques the papers use, named apart from their topics
microbenchmarks · 0.2application benchmarks · 0.2performance tuning · 0.1numerical algorithm evaluation · 0.1performance assertions · 0.1vectorization · 0.1lattice boltzmann method · 0.1data decomposition · 0.1microbenchmarking · 0.0application code evaluation · 0.0runtime monitoring · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Communication Characterization and Optimization of Applications Using Topology-Aware Task Mapping on Large SupercomputersabstractOn large supercomputers, the job scheduling systems may assign a non-contiguous node allocation for user applications depending on available resources. With parallel applications using MPI (Message Passing Interface), the default process ordering does not take into account the actual physical node layout available to the application. This contributes to non-locality in terms of physical network topology and impacts communication performance of the application. In order to mitigate such performance penalties, this work describes techniques to identify suitable task mapping that takes the layout of the allocated nodes as well as the application's communication behavior into account. During the first phase of this research, we instrumented and collected performance data to characterize communication behavior of critical US DOE (United States - Department of Energy) applications using an augmented version of the mpiP tool. Subsequently, we developed several reordering methods (spectral bisection, neighbor join tree etc.) to combine node layout and application communication data for optimized task placement. We developed a tool called mpiAproxy to facilitate detailed evaluation of the various reordering algorithms without requiring full application executions. This work presents a comprehensive performance evaluation (14,000 experiments) of the various task mapping techniques in lowering communication costs on Titan, the leadership class supercomputer at Oak Ridge National Laboratory. Sarat Sreepathi, Eduardo F. D'Azevedo, Bobby Philip, Patrick H. Worley |
ICPE | 4 |
| 2011 | Performance of the community earth system modelabstractThe Community Earth System Model (CESM), released in June 2010, incorporates new physical process and new numerical algorithm options, significantly enhancing simulation capabilities over its predecessor, the June 2004 release of the Community Climate System Model. CESM also includes enhanced performance tuning options and performance portability capabilities. This paper describes performance and performance scaling on both the Cray XT5 and the IBM BG/P for four representative production simulations, varying both problem size and enabled physical processes. The paper also describes preliminary performance results for high resolution simulations using over 200,000 processor cores, indicating the promise of ongoing work in numerical algorithms and where further work is required. Patrick H. Worley, Arthur A. Mirin, Anthony P. Craig, Mark A. Taylor, John M. Dennis, Mariana Vertenstein |
SC | 1 |
| 2008 | Early evaluation of IBM BlueGene/PabstractBlueGene/P (BG/P) is the second generation BlueGene architecture from IBM, succeeding BlueGene/L (BG/L). BG/P is a system-on-a-chip (SoC) design that uses four PowerPC 450 cores operating at 850 MHz with a double precision, dual pipe floating point unit per core. These chips are connected with multiple interconnection networks including a 3-D torus, a global collective network, and a global barrier network. The design is intended to provide a highly scalable, physically dense system with relatively low power requirements per flop. In this paper, we report on our examination of BG/P, presented in the context of a set of important scientific applications, and as compared to other major large scale supercomputers in use today. Our investigation confirms that BG/P has good scalability with an expected lower performance per processor when compared to the Cray XT4's Opteron. We also find that BG/P uses very low power per floating point operation for certain kernels, yet it has less of a power advantage when considering science driven metrics for mission applications. Sadaf R. Alam, Richard F. Barrett, M. Bast, Mark R. Fahey, Jeffery A. Kuehn, Collin McCurdy, James H. Rogers, Philip C. Roth, Ramanan Sankaran, Jeffrey S. Vetter, Patrick H. Worley, Weikuan Yu |
SC | 11 |
| 2008 | Algorithm 888: Spherical Harmonic Transform AlgorithmsabstractA collection of MATLAB classes for computing and using spherical harmonic transforms is presented. Methods of these classes compute differential operators on the sphere and are used to solve simple partial differential equations in a spherical geometry. The spectral synthesis and analysis algorithms using fast Fourier transforms and Legendre transforms with the associated Legendre functions are presented in detail. A set of methods associated with a spectral_field class provides spectral approximation to the differential operators ∇ ⋯, ∇ ×, ∇, and ∇ 2 in spherical geometry. Laplace inversion and Helmholtz equation solvers are also methods for this class. The use of the class and methods in MATLAB is demonstrated by the solution of the barotropic vorticity equation on the sphere. A survey of alternative algorithms is given and implementations for parallel high performance computers are discussed in the context of global climate and weather models. John B. Drake, Patrick H. Worley, Eduardo F. D'Azevedo |
ACM Trans. Math. Softw. | 2 |
| 2007 | Cray XT4: an early evaluation for petascale scientific simulationabstractThe scientific simulation capabilities of next generation high-end computing technology will depend on striking a balance among memory, processor, I/O, and local and global network performance across the breadth of the scientific simulation space. The Cray XT4 combines commodity AMD dual core Opteron processor technology with the second generation of Cray's custom communication accelerator in a system design whose balance is claimed to be driven by the demands of scientific simulation. This paper presents an evaluation of the Cray XT4 using micro-benchmarks to develop a controlled understanding of individual system components, providing the context for analyzing and comprehending the performance of several petascale-ready applications. Results gathered from several strategic application domains are compared with observations on the previous generation Cray XT3 and other high-end computing systems, demonstrating performance improvements across a wide variety of application benchmark problems. Sadaf R. Alam, Jeffery A. Kuehn, Richard F. Barrett, Jeffrey M. Larkin, Mark R. Fahey, Ramanan Sankaran, Patrick H. Worley |
SC | 7 |
| 2006 | Early evaluation of the Cray XT3abstractOak Ridge National Laboratory recently received delivery of a 5,294 processor Cray XT3. The XT3 is Cray's third-generation massively parallel processing system. The system builds on a single processor node - built around the AMD Opteron - and uses a custom chip - called SeaStar - to provide interprocess or communication. In addition, the system uses a lightweight operating system on the compute nodes. This paper describes our initial experiences with the system, including micro-benchmark, kernel, and application benchmark results. In particular, we provide performance results for strategic Department of Energy applications areas including climate and fusion. We demonstrate experiments on the installed system, scaling applications up to 4,096 processors. Jeffrey S. Vetter, Sadaf R. Alam, Thomas H. Dunigan, Mark R. Fahey, Philip C. Roth, Patrick H. Worley |
IPDPS | 6 |
| 2005 | Performance Evaluation of the SGI Altix 3700abstractSGI recently introduced the Altix 3700. In contrast to previous SGI systems, the Altix uses a modified version of the open source Linux operating system and the latest Intel IA-64 processors, the Intel Itanium2. The Altix also uses the next generation SGI interconnect, Numalink3 and NUMAflex, which provides a NUMA, cache-coherent, shared memory, multi-processor system. In this paper, we present a performance evaluation of the SGI Altix using microbenchmarks, kernels, and mission applications. We find that the Altix provides many advantages over other non-vector machines and it is competitive with the Cray XI on a number of kernels and applications. The Altix also shows good scaling, and its globally shared memory allows users convenient parallelization with OpenMP or pthreads. Thomas H. Dunigan, Jeffrey S. Vetter, Patrick H. Worley |
ICPP | 3 |
| 2005 | Leading Computational Methods on Scalar and Vector HEC PlatformsabstractThe last decade has witnessed a rapid proliferation of superscalar cache-based microprocessors to build high-end computing (HEC) platforms, primarily because of their generality, scalability, and cost effectiveness. However, the growing gap between sustained and peak performance for full-scale scientific applications on conventional supercomputers has become a major concern in high performance computing, requiring significantly larger systems and application scalability than implied by peak performance in order to achieve desired performance. The latest generation of custom-built parallel vector systems have the potential to address this issue for numerical algorithms with sufficient regularity in their computational structure. In this work we explore applications drawn from four areas: atmospheric modeling (CAM), magnetic fusion (GTC), plasma physics (LBMHD3D), and material science (PARATEC). We compare performance of the vector-based Cray X1, Earth Simulator, and newly-released NEC SX-8 and Cray X1E, with performance of three leading commodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors. Our work makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new data-decomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES promodity-based superscalar platforms utilizing the IBM Power3, Intel Itanium2, and AMD Opteron processors, with modern parallel vector systems: the Cray X1, Earth Simulator (ES), and the NEC SX-8. Additionally, we examine performance of CAM on the recently-released Cray X1E. Our research team was the first international group to conduct a performance evaluation study at the Earth Simulator Center; remote ES access is not available. Our work builds on our previous efforts [16, 17] and makes several significant contributions: the first reported vector performance results for CAM simulations utilizing a finite-volume dynamical core on a high-resolution atmospheric grid; a new datadecomposition scheme for GTC that (for the first time) enables a breakthrough of the Teraflop barrier; the introduction of a new three-dimensional Lattice Boltzmann magneto-hydrodynamic implementation used to study the onset evolution of plasma turbulence that achieves over 26Tflop/s on 4800 ES processors; and the largest PARATEC cell size atomistic simulation to date. Overall, results show that the vector architectures attain unprecedented aggregate performance across our application suite, demonstrating the tremendous potential of modern parallel vector systems. Leonid Oliker, Jonathan Carter 0002, Michael F. Wehner, Andrew Canning, Stéphane Ethier, Arthur A. Mirin, David Parks, Patrick H. Worley, Shigemune Kitawaki, Yoshinori Tsuda |
SC | 8 |
| 2005 | Practical performance portability in the Parallel Ocean Program (POP)abstractAbstract The design of the Parallel Ocean Program (POP) is described with an emphasis on portability. Performance of POP is presented on a wide variety of computational architectures, including vector architectures and commodity clusters. Analysis of POP performance across machines is used to characterize performance and identify improvements while maintaining portability. A new design of the POP model, including a cache blocking and land point elimination scheme, is described with some preliminary performance results. Published in 2005 by John Wiley & Sons, Ltd. Philip W. Jones, Patrick H. Worley, Yoshikatsu Yoshida, James B. White III, John M. Levesque |
Concurr. Pract. Exp. | 2 |
| 2003 | Early Evaluation of the Cray X1abstractOak Ridge National Laboratory installed a 32 processor Cray X1 in March, 2003, and will have a 256 processor system installed by October, 2003. In this paper we describe our initial evaluation of the X1 architecture, focusing on microbenchmarks, kernels, and application codes that highlight the performance characteristics of the X1 architecture and indicate how to use the system most efficiently. Thomas H. Dunigan, Mark R. Fahey, James B. White III, Patrick H. Worley |
SC | 4 |
| 2002 | Asserting performance expectationsabstractTraditional techniques for performance analysis provide a means for extracting and analyzing raw performance information from applications. Users then compare this raw data to their performance expectations for application constructs. This comparison can be tedious for the scale of today's architectures and software systems. To address this situation, we present a methodology and prototype that allows users to assert performance expectations explicitly in their source code using performance assertions. As the application executes, each performance assertion in the application collects data implicitly to verify the assertion. By allowing the user to specify a performance expectation with individual code segments, the runtime system can jettison raw data for measurements that pass their expectation, while reacting to failures with a variety of responses. We present several compelling uses of performance assertions with our operational prototype, including raising a performance exception, validating a performance model, and adapting an algorithm empirically at runtime. Jeffrey S. Vetter, Patrick H. Worley |
SC | 2 |
| 2002 | Scaling the unscalable: a case study on the AlphaServer SCabstractA case study of the optimization of a climate modeling application on the Compaq AlphaServer SC at the Pittsburgh Supercomputer Center is used to illustrate tools and techniques that are important to achieving good performance scaling. Patrick H. Worley |
SC | 1 |
| 2002 | Early evaluation of the IBM p690abstractOak Ridge National Laboratory recently received 27 32-way IBM pSeries 690 SMP nodes. In this paper, we describe our initial evaluation of the p690 architecture, focusing on the performance of benchmarks and applications that are representative of the expected production workload. Patrick H. Worley, Thomas H. Dunigan, Mark R. Fahey, James B. White III, Arthur S. Bland |
SC | 1 |
| 2000 | Performance evaluation of the IBM SP and the Compaq AlphaServer SCabstractOak Ridge National Laboratory (ORNL) has recently installed both a Compaq AlphaServer SC and an IBM SP, each with 4-way SMP nodes, allowing a direct comparison of the two architectures. In this paper, we describe our initial evaluation. The evaluation looks at both kernel and application performance for a spectral atmospheric general circulation model, an important application for the ORNL systems. Patrick H. Worley |
ICS | 1 |
| 1999 | Performance Tuning and Evaluation of a Parallel Community Climate ModelabstractThe Parallel Community Climate Model (PCCM) is a message-passing parallelization of version 2.1 of the Community Climate Model (CCM) developed by researchers at Argonne and Oak Ridge National Laboratories and at the National Center for Atmospheric Research in the early to mid 1990s. In preparation for use in the Department of Energy’s Parallel Climate Model (PCM), PCCM has recently been updated with new physics routines from version 3.2 of the CCM, improvements to the parallel implementation, and ports to the SGI/Cray Research T3E and Origin 2000. We describe our experience in porting and tuning PCCM on these new platforms, evaluating the performance of different parallel algorithm options and comparing performance between the T3E and Origin 2000. 1 John B. Drake, Steve Hammond, Rodney James, Patrick H. Worley |
SC | 4 |
| 1998 | Performance modeling for SPMD message-passing programsabstractToday's massively parallel machines are typically message-passing systems consisting of hundreds or thousands of processors. Implementing parallel applications efficiently in this environment is a challenging task, and poor parallel design decisions can be expensive to correct. Tools and techniques that allow the fast and accurate evaluation of different parallelization strategies would significantly improve the productivity of application developers and increase throughput on parallel architectures. This paper investigates one of the major issues in building tools to compare parallelization strategies: determining what type of performance models of the application code and of the computer system are sufficient for a fast and accurate comparison of different strategies. The paper is built around a case study employing the performance prediction tool (PerPreT) to predict performance of the parallel spectral transform shallow water model code (PSTSWM) on the Intel Paragon. PSTSWM is a parallel application code that was designed to evaluate different parallel strategies for the spectral transform method as it is used in climate modeling and weather forecasting. Multiple parallel algorithms and algorithm variants are embedded in the code. PerPreT uses a relatively simple algebraic model to predict execution time for SPMD (single program multiple data) parallel applications. Applications are modeled through parameterized formulae for communication and computation, where the parameters include the problem size, the number of processors used to execute the program, and system characteristics (e.g. setup times for communication, link bandwidth and sustained computing performance per processor). In this paper we describe performance models that predict the performance of the different algorithms in PSTSWM accurately enough to allow them to be compared, establishing the feasibility of such a demanding application of performance modeling. We also discuss issues in generating and validating the performance models, emphasizing the practical importance of tools such as PerPreT in such studies. © 1998 John Wiley & Sons, Ltd. Jürgen Brehm, Patrick H. Worley, Manish Madhukar |
Concurr. Pract. Exp. | 2 |
| 1998 | A study of application sensitivity to variation in message-passing latency and bandwidthabstractThis study measures the effects of changes in message latency and bandwidth for production-level codes on a current generation tightly coupled MPP, the Intel Paragon. Messages are sent multiple times to study the application sensitivity to variations in bandwidth and latency. This method preserves the effects of contention on the interconnection network. Two applications are studied: PCTH, a shock physics code developed at Sandia National Laboratories; and PSTSWM, a spectral shallow water code developed at Oak Ridge National Laboratory and Argonne National Laboratory. These codes are significant in that PCTH is a ‘full physics’ application code in production use, while PSTSWM serves as a parallel algorithm test bed and benchmark for production codes used in atmospheric modeling. They are also significant in that the message-passing behavior differs significantly between the two codes, each representing an important class of scientific message-passing applications. © 1998 John Wiley & Sons, Ltd. Patrick H. Worley, Allen C. Robinson, David R. Mackay, Edward J. Barragy |
Concurr. Pract. Exp. | 1 |
| 1995 | Design and Performance of a Scalable Parallel Community Climate Model
John B. Drake, Ian T. Foster, John Michalakes, Brian R. Toonen, Patrick H. Worley |
Parallel Comput. | 5 |
| 1992 | Parallelizing the spectral transform method. Part IIabstractAbstract The spectral transform method is a widely used numerical technique for solving partial differential equations on the sphere in global climate modeling. This paper describes the parallelization and performance of the spectral method for solving the non‐linear shallow water equations on the surface of a sphere using a 128‐node Intel iPSC/860 hypercube. Solving the shallow water equations represents a computational kernel of more complex climate models. This work is part of a research program to develop climate models that are capable of much longer simulations at a significantly finer resolution than current models. Such models are important in understanding the effects of the increasing atmospheric concentrations of greenhouse gases, and the computational requirements are so large that massively parallel multiprocessors will be necessary to run climate model simulations in a reasonable amount of time. The spectral method involves the transformation of data between the physical, Fourier and spectral domains. Each of these domains is two‐dimensional. The spectral method performs Fourier transforms in the longitude direction followed by summation in the latitude direction to evaluate the discrete spectral transform. A simple way of parallelizing the spectral code is to decompose the physical problem domain in just the latitude direction. This allows an optimized sequential FFT algorithm to be used in the longitude direction. However, this approach limits the number of processors that can be brought to bear on the problem. Decomposing the problem over both directions allows the parallelism inherent in the problem to be exploited more effectively‐the grain size is reduced, so that more processors can be used. Results are presented that show that decomposing over both directions does result in a more rapid solution of the problem. The results show that, for a given problem and number of processors, the optimum decomposition has approximately equal numbers of processors in each direction. Load imbalance also has an impact on the performance of the method. The importance of minimizing communication latency and overlapping communication with calculation is stressed. General methods for doing this, that may be applied to many other problems, are discussed. David W. Walker, Patrick H. Worley, John B. Drake |
Concurr. Pract. Exp. | 2 |
| 1992 | Parallelizing the spectral transform methodabstractAbstract The spectral transform method is a standard numerical technique used to solve partial differential equations on the sphere in global climate modeling. In particular, it is used in CCM1 and CCM2, the Community Climate Models developed at the National Center for Atmospheric Research. This paper describes initial experiences in parallelizing a program that uses the spectral transform method to solve the non‐linear shallow water equations on the sphere, showing that an efficient implementation is possible on the Intel iPSC/860. The use of PICL, a portable instrumented communication library, and Paragraph, a performance visualization tool, in tuning the implementation is also described. The Legendre transform and the Fourier transform comprise the computational kernel of the spectral transform method. This paper is a case study of parallelizing the Legendre transform. For many problem sizes and numbers of processors, the spectral transform method can be parallelized efficiently by parallelizing only the Legendre transform. Patrick H. Worley, John B. Drake |
Concurr. Pract. Exp. | 1 |
| 1992 | The effect of multiprocessor radius on scaling
Patrick H. Worley |
Parallel Comput. | 1 |