Stephen W. Poole

dblp:48/3377 · also Stephen Poole 0001, Steve Poole 0001 · DBLP profile ↗
← Back
35ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-4531-7453ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 6 since 2021Software engineering, systems software and programming languages · 3Computer networks · 2Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Performance Analysis of Conveyors: Memory Dominates?
abstract
Small-message aggregation is critical for scaling irregular, communication intensive applications in high-performance computing. In this paper, contrary to conventional wisdom, we present the first systematic study showing that memory contention, not network bandwidth, is the dominant bottleneck in message aggregation runtimes. Using the state-of-the-art conveyors library as our reference implementation, we conducted extensive experiments on HPC systems featuring Slingshot 11 and InfiniBand interconnects, scaling to 16k cores (256 nodes) and processing 10s–100s GB of data. Our measurements reveal that interference between user data and aggregation buffers drives LLC miss rates to 77%, inflating memory costs by 2–3× over the algorithmic baseline. Consequently, we advocate for dedicated near-memory subsystems to improve the scalability and performance of message aggregation runtimes. This paper also demonstrates up to an order of magnitude higher latency for conveyor termination compared to a traditional HPC barrier, and it examines the impact of communication context isolation and the critical challenge of programmability.
Shubhendra Pal Singhal, Aaron Welch, Oscar R. Hernandez, Stephen W. Poole, Akihiro Hayashi, Vivek Sarkar
HPDC4
2024 Effective and Efficient Offloading Designs for One-Sided Communication to SmartNICs
abstract
One-sided communication is one of many approaches to use for data transfer in High-Performance Computing (HPC) applications. One-sided operations require less demand on parallel programming libraries and do not require HPC hardware to issue acknowledgments of successful data transfer. Thanks to its inherently non-blocking nature, one-sided communication is also useful for improving overlap between communication and compute. As with any non-blocking communication, however, we run into the issue of message progression getting interleaved with computation. With the advent of Smart Network Cards (SmartNIC) such as NVIDIA's BlueField Data Processing Units (DPU), we can offload the communication and message progression to these devices to improve the overlap of communication and compute. In this paper, we propose designs for efficient offloading of one-sided communication. We show how our designs can be used for offloading both MPI one-sided “put” and “get” and OpenSHMEM's non-blocking “put” and “get”. Using a Block Sparse Matrix-Multiplication Kernel (BSPMM), we show that our designs achieve over 96% improvement in runtime over pure-host execution for communication offload. We also briefly explore initial compute offload ideas for such one-sided kernels and show over 91% improvement in runtime here.
Benjamin Michalowicz, Kaushik Kandadi Suresh, Hari Subramoni, Mustafa Abdul Jabbar, Dhabaleswar K. Panda 0001, Stephen W. Poole
HiPC6
2023 Extending OpenSHMEM with Aggregation Support for Improved Message Rate Performance
Aaron Welch, Oscar R. Hernandez, Stephen W. Poole
Euro-Par3
2023 Battle of the BlueFields: An In-Depth Comparison of the BlueField-2 and BlueField-3 SmartNICs
abstract
Over the past several years, Smart Network Interface Cards (NIC/SmartNICs) have rapidly evolved in popularity. In particular, NVIDIA’s BlueField line of SmartNICs has been effective in a wide variety of uses: Offloading communication in High-Performance Computing applications (HPC), various stages of the Deep Learning (DL) pipeline, and is designed especially for Datacenter/virtualization uses. The BlueField-3 DPU was released at the end of 2022 as a follow-up to its widely accepted BlueField-2 predecessor, and this work will serve as an in-depth performance evaluation between the two to show a) a comparison of both SmartNICs’ on-chip capabilities (memory bandwidth, compute speed, etc.), and b) their offload capabilities through several micro/benchmarks and applications. In single-DPU programs, we see up to 61% improvements in the latency of a memcpy operation and up to 82% bandwidth improvement in the use of the STREAM benchmark [8] on the BlueField-3. With the use of a DPU-aware MPI library [1], we observe over 30% improvement at the micro-benchmark level when comparing staging-based designs on both SmartNICs and up to nearly double that in the context of an application with staging-based designs. However, GVMI (Guest Virtual Machine ID) based designs contained in said library do not exceed 10% at the benchmark level and provide less than 2% benefits in applications because of its architecture-insensitive nature — that is, while CPU clock speed may impact the completion time of instructions, the performance of the GVMI-based designs in a DPU-aware MPI library will largely be unaffected by swapping the BlueField-2 for a BlueField-3.
Benjamin Michalowicz, Kaushik Kandadi Suresh, Hari Subramoni, Dhabaleswar K. Panda 0001, Stephen W. Poole
HOTI5
2022 Bring the BitCODE-Moving Compute and Data in Distributed Heterogeneous Systems
abstract
In this paper, we present a framework for moving compute and data between processing elements in a distributed heterogeneous system. The implementation of the framework is based on the LLVM compiler toolchain combined with the UCX communication framework. The framework can generate binary machine code or LLVM bitcode for multiple CPU architectures and move the code to remote machines while dynamically optimizing and linking the code on the target platform. The remotely injected code can recursively propagate itself to other remote machines or generate new code. The goal of this paper is threefold: (a) to present an ar-chitecture and implementation of the framework that provides essential infrastructure to program a new class of disaggregated systems wherein heterogeneous programming elements such as compute nodes and data processing units (DPUs) are distributed across the system, (b) to demonstrate how the framework can be integrated with modern, high-level programming languages such as Julia, and (c) to demonstrate and evaluate a new class of eXtended Remote Direct Memory Access (X-RDMA) communication operations that are enabled by this framework. To evaluate the capabilities of the framework, we used a cluster with Fujitsu CPUs and heterogeneous cluster with Intel CPUs and BlueField-2 DPUs interconnected using high-performance RDMA fabric. We demonstrated an X-RDMA pointer chase application that outperforms an RDMA GET-based implementation by 70% and is as fast as Active Messages, but does not require function predeployment on remote platforms.
Wenbin Lu, Luis E. Peña, Pavel Shamis, Valentin Churavy, Barbara M. Chapman, Stephen W. Poole
CLUSTER6
2021 Two-Chains: High Performance Framework for Function Injection and Execution
Megan Grodowitz, Luis E. Peña, Curtis Dunham, Dong Zhong, Pavel Shamis, Stephen W. Poole
CLUSTER6
2017 Thoughtful Precision in Mini-Apps
abstract
Approximate computing addresses many of the identified challenges for exascale computing, leading to performance improvements that may include changes in fidelity of calculation. In this paper, we examine approximate approaches for a range of DOE-relevant computational problems run on a variety of architectures as a proxy for the wider set of exascaleclass applications.We show anticipated improvements in computational and memory performance and in power savings. We also assess application correctness when operating under conditions of reduced precision, and show that this is within acceptable bounds. Finally, we discuss the trade space between performance, power, precision and resolution for these mini-apps, and optimized solutions attained within given constraints, with positive implications for application of approximate computing to exascale-class problems.
Shane Fogerty, Siddhartha Bishnu, Yuliana Zamora, Laura Monroe, Stephen W. Poole, Michael O. Lam, Joe Schoonover, Robert W. Robey
CLUSTER5
2015 Measuring Server Energy Proportionality
abstract
In performance engineering, metrics are often used to track the progress over time. Concerning the potential bias of using a single metric, performance engineers tend to use multiple metrics for reasoning. However, this approach has its own challenges. In this work we study one of the challenges in the context of analyzing trends in server energy proportionality. We examine a wide range of metrics for measuring energy proportionality, trying to determine which metrics are essential and which are redundant. We do this by comparing the trend curves of the metrics for the published results of the SPECpower_ssj2008 benchmark. While the context is specific, the proposed analysis method is quite general. We hope that this method would help us do performance engineering more effectively.
Chung-Hsing Hsu, Stephen W. Poole
ICPE2
2015 Utility Functions and Resource Management in an Oversubscribed Heterogeneous Computing Environment
abstract
We model an oversubscribed heterogeneous computing system where tasks arrive dynamically and a scheduler maps the tasks to machines for execution. The environment and workloads are based on those being investigated by the Extreme Scale Systems Center at Oak Ridge National Laboratory. Utility functions that are designed based on specifications from the system owner and users are used to create a metric for the performance of resource allocation heuristics. Each task has a time-varying utility (importance) that the enterprise will earn based on when the task successfully completes execution. We design multiple heuristics, which include a technique to drop low utility-earning tasks, to maximize the total utility that can be earned by completing tasks. The heuristics are evaluated using simulation experiments with two levels of oversubscription. The results show the benefit of having fast heuristics that account for the importance of a task and the heterogeneity of the environment when making allocation decisions in an oversubscribed environment. The ability to drop low utility-earning tasks allow the heuristics to tolerate the high oversubscription as well as earn significant utility.
Bhavesh Khemka, Ryan D. Friese, Luis Diego Briceno, Howard Jay Siegel, Anthony A. Maciejewski, Gregory A. Koenig, Chris Groër, Gene Okonski, Marcia Hilton, Jendra Rambharos, Stephen W. Poole
IEEE Trans. Computers11
2014 Power Consumption Due to Data Movement in Distributed Programming Models
Siddhartha Jana, Oscar R. Hernandez, Stephen W. Poole, Barbara M. Chapman
Euro-Par3
2013 Exploring energy and performance behaviors of data-intensive scientific workflows on systems with deep memory hierarchies
abstract
The increasing gap between the rate at which large scale scientific simulations generate data and the corresponding storage speeds and capacities is leading to more complex system architectures with deep memory hierarchies. Advances in non-volatile memory (NVRAM) technology have made it an attractive candidate as intermediate storage in this memory hierarchy to address the latency and performance gap between main memory and disk storage. As a result, it is important to understand and model its energy/performance behavior from an application perspective as well as how it can be effectively used for staging data within an application workflow. In this paper, we target a NVRAM-based deep memory hierarchy and explore its potential for supporting in-situ/in-transit data analytics pipelines that are part of application workflows patterns. Specifically, we model the memory hierarchy and experimentally explore energy/performance behaviors of different data management strategies and data exchange patterns, as well as the tradeoffs associated with data placement, data movement and data processing.
Marc Gamell, Ivan Rodero, Manish Parashar, Stephen W. Poole
HiPC4
2013 Revisiting Server Energy Proportionality
abstract
Server energy proportionality refers to a proposed ideal that a server consumes energy proportional to its utilization level. Since its introduction in 2007, the ideal has been adopted by the server industry as a design goal to further optimize the energy efficiency of their servers. There have also been studies on how energy proportionality has evolved over time. However, most of these efforts assume linear proportionality. Must energy proportionality be linear? In this paper we look into this problem. We analyze 410 SPECpower_ssj2008 benchmark results published from 2007 to 2012, and find that modern servers have pushed server energy proportionality from linear to quadratic. We also observe a strong correlation in time between this change and dynamic over-clocking (such as Intel Turbo Boost). We present all these findings in the paper, along with discussions on the implications of super-linearity in server energy proportionality.
Chung-Hsing Hsu, Stephen W. Poole
ICPP2
2012 Enabling event tracing at leadership-class scale through I/O forwarding middleware
abstract
Event tracing is an important tool for understanding the performance of parallel applications. As concurrency increases in leadership-class computing systems, the quantity of performance log data can overload the parallel file system, perturbing the application being observed. In this work we present a solution for event tracing at leadership scales. We enhance the I/O forwarding system software to aggregate and reorganize log data prior to writing to the storage system, significantly reducing the burden on the underlying file system for this type of traffic. Furthermore, we augment the I/O forwarding system with a write buffering capability to limit the impact of artificial perturbations from log data accesses on traced applications. To validate the approach, we modify the Vampir tracing toolset to take advantage of this new capability and show that the approach increases the maximum traced application size by a factor of 5x to more than 200,000 processes.
Thomas Ilsche, Joseph Schuchart, Jason Cope, Dries Kimpe, Terry R. Jones, Andreas Knüpfer, Kamil Iskra, Robert B. Ross, Wolfgang E. Nagel, Stephen W. Poole
HPDC10
2012 The Network Adapter: The Missing Link between MPI Applications and Network Performance
abstract
Network design aspects that influence cost and performance can be classified according to their distance from the applications, into issues concerning topology, switch technology, link technology, network adapter, and communication library. The network adapter has a privileged position to take decisions with more global information than any other component in the network. It receives feedback from the switches and requests from the communication libraries and applications. Also, compared to a network switch, an adapter has access to significantly more memory (host memory and on-chip memory) and memory bandwidth (which typically exceeds network bandwidth). The potential of the adapter to improve global network performance has not yet been fully exploited. In this work we show a series of noticeable performance improvements (of at least 10% to 15%) for medium-sized message exchanges in typical HPC communication patterns by optimizing message segmentation and packet injection policies, that can be implemented in an adapter's firmware inexpensively. We also show that implementing equivalent solutions in the switch (as opposed to the adapter) leads to only marginal performance improvements as the ones obtained by controlling the segmentation and injection policy at the adapter, while involving significantly more cost. In addition, enhancing the adapter will lead to less hardware complexity in the switches, thus reducing cost and energy consumption.
Cyriel Minkenberg, Ronald P. Luijten, Ramón Beivide, Patrick Geoffray, Jesús Labarta, Mateo Valero, Stephen W. Poole
SBAC-PAD8
2012 Towards efficient supercomputing: searching for the right efficiency metric
abstract
Efficiency in supercomputing has traditionally focused on execution time. In early 2000's, the concept of total cost of ownership was re-introduced, with the introduction of efficiency measure to include aspects such as energy and space. Yet the supercomputing community has never agreed upon a metric that can cover these aspects completely and also provide a fair basis for comparison. This paper examines the metrics that have been proposed in the past decade, and proposes a vector-valued metric for efficient supercomputing. Using this metric, the paper presents a study of where the supercomputing industry has been and where it stands today with respect to efficient supercomputing.
Chung-Hsing Hsu, Jeffery A. Kuehn, Stephen W. Poole
ICPE3
2011 Diagnosing Anomalous Network Performance with Confidence
abstract
Variability in network performance is a major obstacle in effectively analyzing the throughput of modern high performance computer systems. High performance interconnection networks offer excellent best-case network latencies, however, highly parallel applications running on parallel machines typically require consistently high levels of performance to adequately leverage the massive amounts of available computing power. Performance analysts have usually quantified network performance using traditional summary statistics that assume the observational data is sampled from a normal distribution. In our examinations of network performance, we have found this method of analysis often provides too little data to understand anomalous network performance. In particular, we examine a multi-modal performance scenario encountered with an Infiniband interconnection network and we explore the performance repeatability on the custom Cray SeaStar2 interconnection network after a set of software and driver updates.
Bradley W. Settlemyer, Stephen W. Hodson, Jeffery A. Kuehn, Stephen W. Poole
CCGRID4
2011 Reducing Energy Usage with Memory and Computation-Aware Dynamic Frequency Scaling
Michael Laurenzano, Mitesh R. Meswani, Laura Carrington, Allan Snavely, Mustafa M. Tikir, Stephen W. Poole
Euro-Par (1)6
2011 An idiom-finding tool for increasing productivity of accelerators
abstract
Suppose one is considering purchase of a computer equipped with accelerators. Or suppose one has access to such a computer and is considering porting code to take advantage of the accelerators. Is there a reason to suppose the purchase cost or programmer effort will be worth it? It would be nice to able to estimate the expected improvements in advance of paying money or time. We exhibit an analytical framework and tool-set for providing such estimates: the tools first look for user-defined idioms that are patterns of computation and data access identified in advance as possibly being able to benefit from accelerator hardware. A performance model is then applied to estimate how much faster these idioms would be if they were ported and run on the accelerators, and a recommendation is made as to whether or not each idiom is worth the porting effort to put them on the accelerator and an estimate is provided of what the overall application speedup would be if this were done.
Laura Carrington, Mustafa M. Tikir, Catherine Mills Olschanowsky, Michael Laurenzano, Joshua Peraza, Allan Snavely, Stephen W. Poole
ICS7
2011 Power signature analysis of the SPECpower_ssj2008 benchmark
abstract
As the power consumption of a server system becomes a mainstream concern in enterprise environments, understanding the system's power behavior at varying utilization levels provides us a key to select appropriate energy-efficiency optimizations. In this work, we present an in-depth analysis of 177 SPECpower_ssj2008 results published between 2007-2010 to understand the changes of server's power behavior over time. In particular, we identified simple nonlinear functions appropriate for modeling the power behavior of today's, aggressively power-managed, machines. We consider this work as an important first step towards developing capability for power signature analysis of a high-end computer system.
Chung-Hsing Hsu, Stephen W. Poole
ISPASS2
2011 A technique for moving large data sets over high-performance long distance networks
abstract
In this paper we look at the performance characteristics of three tools used to move large data sets over dedicated long distance networking infrastructure. Although performance studies of wide area networks have been a frequent topic of interest, performance analyses have tended to focus on network latency characteristics and peak throughput using network traffic generators. In this study we instead perform an end-to-end long distance networking analysis that includes reading large data sets from a source file system and committing the data to a remote destination file system. An evaluation of end-to-end data movement is also an evaluation of the system configurations employed and the tools used to move the data. For this paper, we have built several storage platforms and connected them with a high performance long distance network configuration. We use these systems to analyze the capabilities of three data movement tools: BBcp, GridFTP, and XDD. Our studies demonstrate that existing data movement tools do not provide efficient performance levels or exercise the storage devices in their highest performance modes.
Bradley W. Settlemyer, Jonathan D. Dobson, Stephen W. Hodson, Jeffery A. Kuehn, Stephen W. Poole, Thomas Ruwart
MSST5
2011 A mathematical analysis of the R-MAT random graph generator
abstract
Abstract The R‐MAT graph generator introduced by Chakrabarti et al (Int Conf Data Mining, 2004) offers a simple, fast method for generating very large directed graphs. These properties have made it a popular choice as a method of generating graphs for objects of study in a variety of disciplines, from social network analysis to high performance computing. We analyze the graphs generated by R‐MAT and model the generator in terms of occupancy problems to prove results about the degree distributions of these graphs. We prove that the limiting degree distributions can be expressed as a mixture of normal distributions with means and variances that can be easily calculated from the R‐MAT parameters. Additionally, this article offers an efficient computational technique for computing the exact degree distribution and concise expressions for a number of properties of R‐MAT graphs. ©2011 Wiley Periodicals, Inc.*. NETWORKS, 2011
Chris Groër, Blair D. Sullivan, Stephen W. Poole
Networks3
2010 ConnectX-2 InfiniBand Management Queues: First Investigation of the New Support for Network Offloaded Collective Operations
abstract
This paper introduces the newly developed Infini-Band (IB) Management Queue capability, used by the Host Channel Adapter (HCA) to manage network task data flow dependancies, and progress the communications associated with such flows. These tasks include sends, receives, and the newly supported wait task, and are scheduled by the HCA based on a data dependency description provided by the user. This functionality is supported by the ConnectX-2 HCA, and provides the means for delegating collective communication management and progress to the HCA, also known as collective communication offload. This provides a means for overlapping collective communications managed by the HCA and computation on the Central Processing Unit (CPU), thus making it possible to reduce the impact of system noise on parallel applications using collective operations. This paper further describes how this new capability can be used to implement scalable Message Passing Interface (MPI) collective operations, describing the high level details of how this new capability is used to implement the MPI_Barrier collective operation, focusing on the latency sensitive performance aspects of this new capability. This paper concludes with small scale benchmark experiments comparing implementations of the barrier collective operation, using the new network offload capabilities, with established point-to-point based implementations of these same algorithms, which manage the data flow using the central processing unit. These early results demonstrate the promise this new capability provides to improve the scalability of high-performance applications using collective communications. The latency of the HCA based implementation of the barrier is similar to that of the best performing point-to-point based implementation managed by the central processing unit, starting to outperform these as the number of processes involved in the collective operation increases.
Richard L. Graham, Stephen W. Poole, Pavel Shamis, Gil Bloch, Noam Bloch, Hillel Chapman, Michael Kagan, Ariel Shahar, Ishai Rabinovitz, Gilad Shainer
CCGRID2
2010 Investigating the potential of application-centric aggressive power management for HPC workloads
abstract
Energy efficiency of large-scale data centers is becoming a major concern not only for reasons of energy conservation, failures, and cost reduction, but also because such sys tems are soon reaching the limits of power available to them. Like High Performance Computing (HPC) systems, large-scale clu ster-based data centers can consume power in megawatts, and of all the power consumed by such a system, only a fraction is used for actual computations. In this paper, we study the potential of application-centric aggressive power management of data center's resources for HPC workloads. Specifically, we consider power management mechanisms and controls (currently or soon to be) available at different levels and for different subsystems, and leverage several innovative approaches that have been taken to tackle this problem in the last few years, can be effectively used in a application-aware manner for HPC workloads. To do this, we first profile sta ndard HPC benchmarks with respect to behaviors, resource usage and power impact on individual computing nodes. Based on a power and latency model and the workload profiles, we develop an algorithm that can improve energy efficiency with little or no performance loss. We then evaluate our proposed algorithm through simulations using empirical power characterization and quantification. Finally, we validate the simulation results with actual executions on real hardware. The obtained results show that by using application aware power management, we can re-du ce the average energy consumption without significant penalty in performance. This motivates us to investigate autonomic approaches for application-aware aggressive power management and cross layer and cross function predictive subsystem level power management for large-scale data centers.
Ivan Rodero, Sharat Chandra, Manish Parashar, Rajeev Muralidhar, Harinarayanan Seshadri, Stephen W. Poole
HiPC6
2010 Sparse Matrix-Vector Multiplication on a Reconfigurable Supercomputer with Application
abstract
Double precision floating point Sparse Matrix-Vector Multiplication (SMVM) is a critical computational kernel used in iterative solvers for systems of sparse linear equations. The poor data locality exhibited by sparse matrices along with the high memory bandwidth requirements of SMVM result in poor performance on general purpose processors. Field Programmable Gate Arrays (FPGAs) offer a possible alternative with their customizable and application-targeted memory sub-system and processing elements. In this work we investigate two separate implementations of the SMVM on an SRC-6 MAPStation workstation. The first implementation investigates the peak performance capability, while the second implementation balances the amount of instantiated logic with the available sustained bandwidth of the FPGA subsystem. Both implementations yield the same sustained performance with the second producing a much more efficient solution. The metrics of processor and application balance are introduced to help provide some insight into the efficiencies of the FPGA and CPU based solutions explicitly showing the tight coupling of the available bandwidth to peak floating point performance. Due to the FPGAs ability to balance the amount of implemented logic to the available memory bandwidth it can provide a much more efficient solution. Finally, making use of the lessons learned implementing the SMVM, we present a fully implemented non-preconditioned Conjugate Gradient Algorithm utilizing the second SMVM design.
David DuBois, Andrew DuBois, Thomas Boorman, Carolyn Connor Davenport, Stephen W. Poole
ACM Trans. Reconfigurable Technol. Syst.5
2009 Impact of Quad-Core Cray XT4 System and Software Stack on Scientific Computation
Sadaf R. Alam, Richard F. Barrett, Heike Jagode, Jeffery A. Kuehn, Stephen W. Poole, Ramanan Sankaran
Euro-Par5
2009 Performance Characterization of a Hierarchical MPI Implementation on Large-scale Distributed-memory Platforms
abstract
The building blocks of emerging Petascale massively parallel processing (MPP) systems are multi-core processors with four or more cores as a single processing element and a customized network interface. The resulting memory and communication hierarchy of these platforms are now exposed to application developers and end users by creating a hierarchical or multi-core aware message-passing (MPI) programming interface and by providing a handful of runtime, tunable parameters that allows mapping and control of MPI tasks and message handling. We characterize performance of MPI communication patterns and present strategies for optimizing applications performance on Cray XT series systems that are composed of contemporary AMD processors and a proprietary network infrastructure. We highlight dependencies in its memory and network subsystems, which could influence production-level applications performance. We demonstrate that MPI micro-benchmarks could mislead an application developer or end user since these benchmarks often do not expose the interplay between memory allocation and usage in the user space, which depends on the number of tasks or cores and workload characteristics. Our studies show performance improvements compared to the default options for our target scientific benchmarks and production-level applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
ICPP4
2009 Performance analysis and projections for Petascale applications on Cray XT series systems
abstract
The Petascale Cray XT5 system at the Oak Ridge National Laboratory (ORNL) Leadership Computing Facility (LCF) shares a number of system and software features with its predecessor, the Cray XT4 system including the quad-core AMD processor and a multi-core aware MPI library. We analyze performance of scalable scientific applications on the quad-core Cray XT4 system as part of the early system access using a combination of micro-benchmarks and Petascale ready applications. Particularly, we evaluate impact of key changes that occurred during the dual-core to quad-core processor upgrade on applications behavior and provide projections for the next-generation massively-parallel platforms with multi-core processors, specifically for proposed Petascale Cray XT5 system. We compare and contrast the quad-core XT4 system features with the upcoming XT5 system and discuss strategies for improving scaling and performance for our target applications.
Sadaf R. Alam, Richard F. Barrett, Jeffery A. Kuehn, Stephen W. Poole
IPDPS4
2008 An Implementation of the Conjugate Gradient Algorithm on FPGAs
abstract
The conjugate gradient is a prominent iterative method for solving systems of sparse linear equations. Large-scale scientific applications often utilize a conjugate gradient solver at their computational core. Since a single iteration of a conjugate gradient solver requires a sparse matrix-vector multiply operation it is imperative that this operation be computed efficiently. In this paper we present a field programmable gate array (FPGA) based implementation of a double precision, non-preconditioned, conjugate gradient solver for finite-element or finite-difference methods. We show that our FPGA implementation can outperform current generation processors while running at a ~30X slower clock rate. Our work utilizes the SRC Computers, Inc. MAPStation hardware platform along with the "Carte" software programming environment.
David DuBois, Andrew DuBois, Thomas Boorman, Carolyn Connor Davenport, Stephen W. Poole
FCCM5
2008 Sparse Matrix-Vector Multiplication on a Reconfigurable Supercomputer
abstract
Double precision floating point Sparse Matrix-Vector Multiplication (SMVM) is a critical computational kernel used in iterative solvers for systems of sparse linear equations. The poor data locality exhibited by sparse matrices along with the high memory bandwidth requirements of SMVM result in poor performance on general purpose processors. Field Programmable Gate Arrays (FPGAs) offer a possible alternative with their customizable and application-targeted memory sub-system and processing elements.
David DuBois, Andrew DuBois, Carolyn Connor Davenport, Stephen W. Poole
FCCM4
2008 Wide-area performance profiling of 10GigE and InfiniBand technologies
abstract
For wide-area high-performance applications, light-paths provide 10Gbps connectivity, and multi-core hosts with PCI-Express can drive such data rates. However, sustaining such end-to-end application throughputs across connections of thousands of miles remains challenging, and the current performance studies of such solutions are very limited. We present an experimental study of two solutions to achieve such throughputs based on: (a) 10Gbps Ethernet with TCP/IP transport protocols, and (b) InfiniBand and its wide-area extensions. For both, we generate performance profiles over 10Gbps connections of lengths up to 8600 miles, and discuss the components, complexity, and limitations of sustaining such throughputs, using different connections and host configurations. Our results indicate that IB solution is better suited for applications with a single large flow, and 10GigE solution is better for those with multiple competing flows.
Nageswara S. V. Rao, Weikuan Yu, William R. Wing, Stephen W. Poole, Jeffrey S. Vetter
SC4
2007 NPU-Based Image Compositing in a Distributed Visualization System
abstract
This paper describes the first use of a Network Processing Unit (NPU) to perform hardware-based image composition in a distributed rendering system. The image composition step is a notorious bottleneck in a clustered rendering system. Furthermore, image compositing algorithms do not necessarily scale as data size and number of nodes increase. Previous researchers have addressed the composition problem via software and/or custom-built hardware. We used the heterogeneous multicore computation architecture of the Intel IXP28XX NPU, a fully programmable commercial off-the-shelf (COTS) technology, to perform the image composition step. With this design, we have attained a nearly four-times performance increase over traditional software-based compositing methods, achieving sustained compositing rates of 22-28 fps on a 1,024 x 1,024 image. This system is fully scalable with a negligible penalty in frame rate, is entirely COTS, and is flexible with regard to operating system, rendering software, graphics cards, and node architecture. The NPU-based compositor has the additional advantage of being a modular compositing component that is eminently suitable for integration into existing distributed software visualization packages.
David Pugmire, Laura Monroe, Carolyn Connor Davenport, Andrew DuBois, David DuBois, Stephen W. Poole
IEEE Trans. Vis. Comput. Graph.6
2006 PaScal - a new parallel and scalable server IO networking infrastructure for supporting global storage/file systems in large-size Linux clusters
abstract
This paper presents the design and implementation of a new I/O networking infrastructure, named PaScal (parallel and scalable I/O networking framework). PaScal is used to support high data bandwidth IP based global storage systems for large scale Linux clusters. PaScal has several unique properties. It employs (1) Multi-level switch-fabric interconnection network by combining high speed interconnects for computing inter-process communication (IPC) requirements and low-cost Gigabit Ethernet interconnect for global IP based storage/file access, (2) A bandwidth on demand scaling I/O networking architecture, (3) open-standard IP networks (routing and switching), (4) multipath routing for load balancing and failover, (5) open shortest path first (OSPF) routing software, and (6) Supporting a global file system in multi-cluster and multi-platform environments. We describe both the hardware and software components of our proposed PaScal. We have implemented the PaScal I/O infrastructure on several large-size Linux clusters at LANL. We have conducted a sequence of parallel MPI-IO assessment benchmarks on LANL's Pink 1024 node Linux cluster and the Panasas global parallel file system. Performance results from our parallel MPI-IO benchmarks on the Pink cluster demonstrate that the PaScal I/O Infrastructure is robust and capable of scaling in bandwidth on large-size Linux clusters
Gary Grider, Hsing-bung Chen, James Nunez, Stephen W. Poole, Rosie Wacha, Parks Fields, Robert Martinez, Paul Martinez, Satsangat Khalsa, Abbie Matthews, Garth A. Gibson
IPCCC4
2002 Granidt: Towards Gigabit Rate Network Intrusion Detection Technology
Maya B. Gokhale, Dave Dubois, Andy Dubois, Mike Boorman, Stephen W. Poole, Vic Hogsett
FPL5
1991 Wide format floating-point math libraries
abstract
Same accuracy in vector mode
Vicky Markstein, Peter W. Markstein, Stephen W. Poole
SC4
1988 Block-iterative finite element computations for incompressible flow problems
abstract
A block-iterative finite element procedure is presented for two-dimensional fluid dynamics computations on multiply-connected domains based on the vorticity - stream function formulation of the incompressible Navier-Stokes equations. The difficulty associated with the convection term in the vorticity transport equation is addressed by using a streamline-upwind/Petrov-Galerkin scheme.
T. E. Tezduyar, Roland Glowinski, J. Liou, Stephen W. Poole
ICS5