Kevin T. Pedretti

dblp:91/1284 · also Kevin Pedretti · DBLP profile ↗
← Back
28ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-1261-0178ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 2 first-author · 2 since 2021Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 32% Parallel and multicore computing · 18% Hardware reliability and fault tolerance · 16%
Software engineering, system software, and programming languages
3 papers
Operating systems · 74% Concurrent programming · 14% Runtime systems and virtual machines · 11%

Topics — the 16 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › supercomputing
supercomputer deployment
0.412020
Chronicles of astra: challenges and lessons from the first petascale arm supercomputer · SC 2020
Interconnection networks and networks-on-chip
network congestion
0.412019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
Parallel and multicore computing
task allocation
0.412019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
Operating systems › kernel
lightweight kernel
0.212015
Achieving Performance Isolation with Lightweight Co-Kernels · HPDC 2015
Cloud and datacenter computing
performance isolation
0.212015
Achieving Performance Isolation with Lightweight Co-Kernels · HPDC 2015
Processor architecture and microarchitecture › instruction set architecture › RISC
ARM architecture
0.212022
Understanding Memory Failures on a Petascale Arm System · HPDC 2022
High-performance computing › supercomputing
exascale systems
0.112011
Evaluating the viability of process replication reliability for exascale systems · SC 2011
Distributed systems
fault tolerance
0.112011
Evaluating the viability of process replication reliability for exascale systems · SC 2011
Distributed systems
replication
0.112011
Evaluating the viability of process replication reliability for exascale systems · SC 2011
Distributed systems › replication
state machine replication
0.112011
Evaluating the viability of process replication reliability for exascale systems · SC 2011
Parallel and multicore computing › parallel programming models › message passing
MPI applications
0.112019
Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019
Operating systems › resource management
memory management
0.112008
SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008
Concurrent programming
shared memory
0.112008
SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008
Parallel and multicore computing
MPI
0.112008
SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008
Parallel and multicore computing
parallel programming models
0.112008
SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008
Memory systems › memory management › virtual memory
virtual memory addressing
0.112008
SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008

Methods — techniques the papers use, named apart from their topics

containerization · 0.9statistical characterization · 0.6field failure data analysis · 0.6virtualization · 0.4co-kernel architecture · 0.4graph partitioning · 0.4geometric partitioning · 0.4simulation · 0.1modeling · 0.1empirical analysis · 0.1lightweight kernel · 0.1fixed offset virtual memory · 0.1
YearPublicationVenuePosition
2024 Characterizing the Impact of Job Execution on the Occurrence of Memory Failures on a Petascale HPC System
abstract
Characterizing the reliability of current and recent high performance (HPC) systems is critical for forecasting how future systems may behave and informing the design of fault tolerance mechanisms. Although research has been conducted to understand memory failures, there are few examples where the occurrence of memory failures is considered in the broader context of system power, temperature, and the execution of user jobs. In this paper, we combine job data with existing power, temperature, and memory failure data collected on a petascale HPC system. By focusing on periods when jobs were running on the system, we identified trends that were not evident in earlier studies of this same data. The inclusion of job data also demonstrated how user behavior can affect the occurrence of memory failures. In conjunction with this paper, we have publicly released the job data used in this paper to complement existing publicly-available data from Astra regarding power, temperature, and the occurrence of correctable memory failures.
Scott Levy, Joshua Hemmert, Kurt B. Ferreira, Kevin T. Pedretti
SBAC-PAD4
2023 Enabling power measurement and control on Astra: The first petascale Arm supercomputer
abstract
Summary Astra, deployed in 2018, was the first petascale supercomputer to utilize processors based on the ARM instruction set. The system was also the first under Sandia's Vanguard program which seeks to provide an evaluation vehicle for novel technologies that with refinement could be utilized in demanding, large‐scale HPC environments. In addition to ARM, several other important first‐of‐a‐kind developments were used in the machine, including new approaches to cooling the datacenter and machine. This article documents our experiences building a power measurement and control infrastructure for Astra. While this is often beyond the control of users today, the accurate measurement, cataloging, and evaluation of power, as our experiences show, is critical to the successful deployment of a large‐scale platform. While such systems exist in part for other architectures, Astra required new development to support the novel Marvell ThunderX2 processor used in compute nodes. In addition to documenting the measurement of power during system bring up and for subsequent on‐going routine use, we present results associated with controlling the power usage of the processor, an area which is becoming of progressively greater interest as data centers and supercomputing sites look to improve compute/energy efficiency and find additional sources for full system optimization.
Ryan E. Grant, Simon D. Hammond, James H. Laros III, Michael J. Levenhagen, Stephen Olivier, Kevin T. Pedretti, Lee Ward, Andrew J. Younge
Concurr. Comput. Pract. Exp.6
2022 Understanding Memory Failures on a Petascale Arm System
abstract
New and novel HPC platforms provide interesting challenges and opportunities. Analysis of these systems can provide a better understanding of both the specific platform being studied as well as large-scale systems in general. Arm is one such architecture that has been explored in HPC for several years, however little is still known about its viability for supporting large-scale production workloads in terms of system reliability. The Astra system at Sandia National Laboratories was the first public peta-FLOPS Arm-based system on the Top500 and has been successfully running production HPC applications for a couple of years. In this paper, we analyze memory failure data collected from Astra while the system was in production running unclassified applications. This analysis revealed several interesting contributions related to both the Arm platform and to HPC systems in general. First, we outline the number of components replaced due to reliability issues in standing-up this first-of-its-kind, large-scale HPC system. We show the distribution differences between correctable DRAM faults and errors on Astra, showing that, not properly accounting for faults can lead to erroneous conclusions. Additionally, we characterize DRAM faults on the system and show contrary to existing work that memory faults are uniformly distributed across CPU socket, DRAM column, bank and rack region, but are not uniform across node, DIMM rank, DIMM slot on the motherboard, and system rack: some racks, ranks and DIMM slots experience more faults than others. Similarly, we show the impact of temperature and power on DRAM correctable errors. Finally, we make a detailed comparison of results presented here with the positional affects found in several previous large-scale reliability studies. The results of this analysis provide valuable guidance to organizations standing-up first-in- class platforms in HPC, organizations using Arm in HPC, and the entire large-scale HPC community in general.
Kurt B. Ferreira, Scott Levy, Joshua Hemmert, Kevin T. Pedretti
HPDC4
2020 Chronicles of astra: challenges and lessons from the first petascale arm supercomputer
abstract
Arm processors have been explored in HPC for several years, however there has not yet been a demonstration of viability for supporting large-scale production workloads. In this paper, we offer a retrospective on the process of bringing up Astra, the first Petascale supercomputer based on 64-bit Arm processors, and validating its ability to run production HPC applications. Through this process several immature technology gaps were addressed, including software stack enablement, Linux bugs at scale, thermal management issues, power management capabilities, and advanced container support. From this experience, several lessons learned are formulated that contributed to the successful deployment of Astra. These insights can be helpful to accelerate deploying and maturing other first-seen HPC technologies. With Astra now supporting many users running a diverse set of production applications at multi-thousand node scales, we believe this constitutes strong supporting evidence that Arm is a viable technology for even the largest-scale supercomputer deployments.
Kevin T. Pedretti, Andrew J. Younge, Simon D. Hammond, James H. Laros III, Matthew L. Curry, Michael J. Aguilar, Robert J. Hoekstra, Ron Brightwell
SC1
2019 Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks
abstract
We present a new method for reducing parallel applications' communication time by mapping their MPI tasks to processors in a way that lowers the distance messages travel and the amount of congestion in the network. Assuming geometric proximity among the tasks is a good approximation of their communication interdependence, we use a geometric partitioning algorithm to order both the tasks and the processors, assigning task parts to the corresponding processor parts. In this way, interdependent tasks are assigned to “nearby” cores in the network. We also present a number of algorithmic optimizations that exploit specific features of the network or application to further improve the quality of the mapping. We specifically address the case of sparse node allocation, where the nodes assigned to a job are not necessarily located in a contiguous block nor within close proximity to each other in the network. However, our methods generalize to contiguous allocations as well, and results are shown for both contiguous and non-contiguous allocations. We show that, for the structured finite difference mini-application MiniGhost, our mapping methods reduced communication time up to 75 percent relative to MiniGhost's default mapping on 128K cores of a Cray XK7 with sparse allocation. For the atmospheric modeling code E3SM/HOMME, our methods reduced communication time up to 31% on 16K cores of an IBM BlueGene/Q with contiguous allocation.
Mehmet Deveci, Karen D. Devine, Kevin T. Pedretti, Mark A. Taylor, Sivasankaran Rajamanickam, Ümit V. Çatalyürek
IEEE Trans. Parallel Distributed Syst.3
2018 Large-Scale System Monitoring Experiences and Recommendations
abstract
Monitoring of High Performance Computing (HPC) platforms is critical to successful operations, can provide insights into performance-impacting conditions, and can inform methodologies for improving science throughput. However, monitoring systems are not generally considered core capabilities in system requirements specifications nor in vendor development strategies. In this paper we present work performed at a number of large-scale HPC sites towards developing monitoring capabilities that fill current gaps in ease of problem identification and root cause discovery. We also present our collective views, based on the experiences presented, on needs and requirements for enabling development by vendors or users of effective sharable end-to-end monitoring capabilities.
Ville Ahlgren, Stefan Andersson, Jim M. Brandt, Nicholas Cardo, Sudheer Chunduri, Jeremy Enos, Parks Fields, Ann C. Gentile, Richard A. Gerber, Michael Gienger, Joe Greenseid, Annette Greiner, Bilel Hadri, Dennis Hoppe, Urpo Kaila, Kaki Kelly, Mark Klein 0002, Alex Kristiansen, Stephen Leak, Mike Mason, Kevin T. Pedretti, Jean-Guillaume Piccinali, Jason Repik, Jim Rogers, Susanna Salminen, Michael T. Showerman, Cary Whitney, Jim Williams
CLUSTER22
2018 Characterizing MPI matching via trace-based simulation
Kurt B. Ferreira, Scott Levy, Kevin T. Pedretti, Ryan E. Grant
Parallel Comput.3
2017 A Tale of Two Systems: Using Containers to Deploy HPC Applications on Supercomputers and Clouds
abstract
Containerization, or OS-level virtualization has taken root within the computing industry. However, container utilization and its impact on performance and functionality within High Performance Computing (HPC) is still relatively undefined. This paper investigates the use of containers with advanced supercomputing and HPC system software. With this, we define a model for parallel MPI application DevOps and deployment using containers to enhance development effort and provide container portability from laptop to clouds or supercomputers. In this endeavor, we extend the use of Sin- gularity containers to a Cray XC-series supercomputer. We use the HPCG and IMB benchmarks to investigate potential points of overhead and scalability with containers on a Cray XC30 testbed system. Furthermore, we also deploy the same containers with Docker on Amazon's Elastic Compute Cloud (EC2), and compare against our Cray supercomputer testbed. Our results indicate that Singularity containers operate at native performance when dynamically linking Cray's MPI libraries on a Cray supercomputer testbed, and that while Amazon EC2 may be useful for initial DevOps and testing, scaling HPC applications better fits supercomputing resources like a Cray.
Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Ron Brightwell
CloudCom2
2017 Enabling Diverse Software Stacks on Supercomputers Using High Performance Virtual Clusters
abstract
While large-scale simulations have been the hallmark of the High Performance Computing (HPC) community for decades, Large Scale Data Analytics (LSDA) workloads are gaining attention within the scientific community not only as a processing component to large HPC simulations, but also as standalone scientific tools for knowledge discovery. With the path towards Exascale, new HPC runtime systems are also emerging in a way that differs from classical distributed computing models. However, system software for such capabilities on the latest extreme-scale DOE supercomputing needs to be enhanced to more appropriately support these types of emerging software ecosystems. In this paper, we propose the use of Virtual Clusters on advanced supercomputing resources to enable systems to support not only HPC workloads, but also emerging big data stacks. Specifically, we have deployed the KVM hypervisor within Cray's Compute Node Linux on a XC-series supercomputer testbed. We also use libvirt and QEMU to manage and provision VMs directly on compute nodes, leveraging Ethernet-over-Aries network emulation. To our knowledge, this is the first known use of KVM on a true MPP supercomputer. We investigate the overhead our solution using HPC benchmarks, both evaluating single-node performance as well as weak scaling of a 32-node virtual cluster. Overall, we find single node performance of our solution using KVM on a Cray is very efficient with near-native performance. However overhead increases by up to 20% as virtual cluster size increases, due to limitations of the Ethernet-over-Aries bridged network. Furthermore, we deploy Apache Spark with large data analysis workloads in a Virtual Cluster, effectively demonstrating how diverse software ecosystems can be supported by High Performance Virtual Clusters.
Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Brian L. Gaines, Ron Brightwell
CLUSTER2
2016 Sweet Spots and Limits for Virtualization
abstract
This year at VEE, we added a panel to discuss the state of virtualization: what problems are solved? what problems are important? and what problems may not be worth solving? The panelist are experts in areas ranging from hardware virtualization up to language-level virtualization.
Carl A. Waldspurger, Emery D. Berger, Abhishek Bhattacharjee, Kevin T. Pedretti, Simon Peter 0001, Christopher J. Rossbach
VEE4
2015 Achieving Performance Isolation with Lightweight Co-Kernels
abstract
Performance isolation is emerging as a requirement for High Performance Computing (HPC) applications, particularly as HPC architectures turn to in situ data processing and application composition techniques to increase system throughput. These approaches require the co-location of disparate workloads on the same compute node, each with different resource and runtime requirements. In this paper we claim that these workloads cannot be effectively managed by a single Operating System/Runtime (OS/R). Therefore, we present Pisces, a system software architecture that enables the co-existence of multiple independent and fully isolated OS/Rs, or enclaves, that can be customized to address the disparate requirements of next generation HPC workloads. Each enclave consists of a specialized lightweight OS co-kernel and runtime, which is capable of independently managing partitions of dynamically assigned hardware resources. Contrary to other co-kernel approaches, in this work we consider performance isolation to be a primary requirement and present a novel co-kernel architecture to achieve this goal. We further present a set of design requirements necessary to ensure performance isolation, including: (1) elimination of cross OS dependencies, (2) internalized management of I/O, (3) limiting cross enclave communication to explicit shared memory channels, and (4) using virtualization techniques to provide missing OS features. The implementation of the Pisces co-kernel architecture is based on the Kitten Lightweight Kernel and Palacios Virtual Machine Monitor, two system software architectures designed specifically for HPC systems. Finally we will show that lightweight isolated co-kernels can provide better performance for HPC applications, and that isolated virtual machines are even capable of outperforming native environments in the presence of competing workloads.
Jiannan Ouyang, Brian Kocoloski, Jack Lange, Kevin T. Pedretti
HPDC4
2014 Demonstrating improved application performance using dynamic monitoring and task mapping
abstract
This work demonstrates the integration of monitoring, analysis, and feedback to perform application-to-resource mapping that adapts to both static architecture features and dynamic resource state. In particular, we present a framework for mapping MPI tasks to compute resources based on run-time analysis of system-wide network data, architecture-specific routing algorithms, and application communication patterns. We address several challenges. Within each node, we collect local utilization data. We consolidate that information to form a global view of system performance, accounting for system-wide factors including competing applications. We provide an interface for applications to query the global information. Then we exploit the system information to change the mapping of tasks to nodes so that system bottlenecks are avoided. We demonstrate the benefit of this monitoring and feedback by remapping MPI tasks based on route-length, bandwidth, and credit-stalls metrics for a parallel sparse matrix-vector multiplication kernel. In the best case, remapping based on dynamic network information in a congested environment recovered 48.9% of the time lost to congestion, reducing matrix-vector multiplication time by 7.8%. Our experiments focus on the Cray XE/XK platform, but the integration concepts are generally applicable to any platform for which applicable metrics and route knowledge can be obtained.
Jim M. Brandt, Karen D. Devine, Ann C. Gentile, Kevin T. Pedretti
CLUSTER4
2014 Exploiting Geometric Partitioning in Task Mapping for Parallel Computers
abstract
We present a new method for mapping applications' MPI tasks to cores of a parallel computer such that communication and execution time are reduced. We consider the case of sparse node allocation within a parallel machine, where the nodes assigned to a job are not necessarily located within a contiguous block nor within close proximity to each other in the network. The goal is to assign tasks to cores so that interdependent tasks are performed by "nearby" cores, thus lowering the distance messages must travel, the amount of congestion in the network, and the overall cost of communication. Our new method applies a geometric partitioning algorithm to both the tasks and the processors, and assigns task parts to the corresponding processor parts. We show that, for the structured finite difference mini-app Mini Ghost, our mapping method reduced execution time 34% on average on 65,536 cores of a Cray XE6. In a molecular dynamics mini-app, Mini MD, our mapping method reduced communication time by 26% on average on 6144 cores. We also compare our mapping with graph-based mappings from the LibTopoMap library and show that our mappings reduced the communication time on average by 15% in MiniGhost and 10% in MiniMD.
Mehmet Deveci, Sivasankaran Rajamanickam, Vitus J. Leung, Kevin T. Pedretti, Stephen Olivier, David P. Bunde, Ümit V. Çatalyürek, Karen D. Devine
IPDPS4
2014 Exascale design space exploration and co-design
Sudip S. Dosanjh, Richard F. Barrett, Douglas Doerfler, Simon D. Hammond, Karl S. Hemmert, Michael A. Heroux, Paul T. Lin, Kevin T. Pedretti, Arun Rodrigues, Timothy G. Trucano, Justin Luitjens
Future Gener. Comput. Syst.8
2012 Application-driven analysis of two generations of capability computing: the transition to multicore processors
abstract
SUMMARY Multicore processors form the basis of most traditional high performance parallel processing architectures. Early experiences with these computers showed significant performance problems, both with regard to computation and inter‐process communication. The transition from Purple, an IBM POWER5‐based machine, to Cielo, a Cray XE6, as the main capability computing platform for the United States Department of Energy's Advanced Simulation and Computing campaign provides an opportunity to reexamine these issues after experiences with a few generations of multicore‐based machines. Experiences with Purple identified some important characteristics that led to strong performance of complex scientific application programs at very large scales. Herein, we compare the performance of some Advanced Simulation and Computing mission critical applications at capability scale across this transition to multicore processors. Copyright © 2012 John Wiley & Sons, Ltd.
Mahesh Rajan, Courtenay T. Vaughan, Douglas Doerfler, Richard F. Barrett, Paul T. Lin, Kevin T. Pedretti, Karl S. Hemmert
Concurr. Comput. Pract. Exp.6
2011 The Impact of Injection Bandwidth Performance on Application Scalability
Kevin T. Pedretti, Ron Brightwell, Douglas Doerfler, Karl S. Hemmert, James H. Laros III
EuroMPI1
2011 Evaluating the viability of process replication reliability for exascale systems
abstract
As high-end computing machines continue to grow in size, issues such as fault tolerance and reliability limit application scalability. Current techniques to ensure progress across faults, like checkpoint-restart, are increasingly problematic at these scales due to excessive overheads predicted to more than double an application's time to solution. Replicated computing techniques, particularly state machine replication, long used in distributed and mission critical systems, have been suggested as an alternative to checkpoint-restart. In this paper, we evaluate the viability of using state machine replication as the primary fault tolerance mechanism for upcoming exascale systems. We use a combination of modeling, empirical analysis, and simulation to study the costs and benefits of this approach in comparison to checkpoint/restart on a wide range of system parameters. These results, which cover different failure distributions, hardware mean time to failures, and I/O bandwidths, show that state machine replication is a potentially useful technique for meeting the fault tolerance demands of HPC applications on future exascale platforms.
Kurt B. Ferreira, Jon Stearley, James H. Laros III, Ron A. Oldfield, Kevin T. Pedretti, Ron Brightwell, Rolf Riesen, Patrick G. Bridges, Dorian C. Arnold
SC5
2011 Minimal-overhead virtualization of a large scale supercomputer
abstract
Virtualization has the potential to dramatically increase the usability and reliability of high performance computing (HPC) systems. However, this potential will remain unrealized unless overheads can be minimized. This is particularly challenging on large scale machines that run carefully crafted HPC OSes supporting tightly-coupled, parallel applications. In this paper, we show how careful use of hardware and VMM features enables the virtualization of a large-scale HPC system, specifically a Cray XT4 machine, with < = 5% overhead on key HPC applications, microbenchmarks, and guests at scales of up to 4096 nodes. We describe three techniques essential for achieving such low overhead: passthrough I/O, workload-sensitive selection of paging mechanisms, and carefully controlled preemption. These techniques are forms of symbiotic virtualization, an approach on which we elaborate.
Jack Lange, Kevin T. Pedretti, Peter A. Dinda, Patrick G. Bridges, Chang Bae, Philip Soltero, Alex Merritt
VEE2
2010 The Impact of System Design Parameters on Application Noise Sensitivity
abstract
Operating system noise, or “jitter,” is a key limiter of application scalability in high end computing systems. Several studies have attempted to quantify the sources and effects of system interference, though few of these studies show the influence that architectural and system characteristics have on the impact of OS noise at scale. In this paper, we examine the impact of three such system properties: platform balance, “noisy” node distribution, and non-blocking collective operations. Using a previouslydeveloped noise injection tool, we explore how the impact of noise varies with these platform characteristics. We provide detailed performance results that indicate that a system with relatively less network bandwidth is able to absorb more noise than a system with more network bandwidth. Our results also show that application performance can be significantly degraded by only a subset of noisy nodes. Furthermore, the placement of the noisy nodes is also important, especially for applications that make substantial use of collective communication operations that are tree-based. Lastly, performance results indicate that nonblocking collective operations have the ability to greatly mitigate the impact of OS interference. Combined, these results show that the impact of OS noise is not solely a property of application communication behavior, but is also influenced by other properties of the system architecture and system software environment.
Kurt B. Ferreira, Patrick G. Bridges, Ron Brightwell, Kevin T. Pedretti
CLUSTER4
2010 Palacios and Kitten: New high performance operating systems for scalable virtualized and native supercomputing
abstract
Palacios is a new open-source VMM under development at Northwestern University and the University of New Mexico that enables applications executing in a virtualized environment to achieve scalable high performance on large machines. Palacios functions as a modularized extension to Kitten, a high performance operating system being developed at Sandia National Laboratories to support large-scale supercomputing applications. Together, Palacios and Kitten provide a thin layer over the hardware to support full-featured virtualized environments alongside Kitten's lightweight native environment. Palacios supports existing, unmodified applications and operating systems by using the hardware virtualization technologies in recent AMD and Intel processors. Additionally, Palacios leverages Kitten's simple memory management scheme to enable low-overhead pass-through of native devices to a virtualized environment. We describe the design, implementation, and integration of Palacios and Kitten. Our benchmarks show that Palacios provides near native (within 5%), scalable performance for virtualized environments running important parallel applications. This new architecture provides an incremental path for applications to use supercomputers, running specialized lightweight host operating systems, that is not significantly performance-compromised.
Jack Lange, Kevin T. Pedretti, Trammell Hudson, Peter A. Dinda, Lei Xia 0001, Patrick G. Bridges, Andy Gocke, Steven Jaconette, Michael J. Levenhagen, Ron Brightwell
IPDPS2
2009 Topics on measuring real power usage on high performance computing platforms
abstract
Power has recently been recognized as one of the major obstacles in fielding a Peta-FLOPs class system. To reach Exa-FLOPs, the challenge will certainly be compounded. In this paper we will discuss a number of High Performance Computing power related topics. We first describe our implementation of a scalable power measurement framework that has enabled us to examine real power use (current draw). [Using this framework, samples were obtained at a per-node (socket) granularity, at frequencies of up to 100 samples per second.] Additionally, we describe how we applied this capability to implement power conserving measures on our Catamount Light Weight Kernel, where we achieved an 80% improvement. This ability has enabled us to quantify the amount of energy used by applications and to contrast application energy use between a Light Weight and General Purpose operating system. Finally, we show application energy use increases proportionally with the increase in run-time due to operating system noise. Areas of future interest will also be discussed.
James H. Laros III, Kevin T. Pedretti, Suzanne M. Kelly, John P. Vandyke, Kurt B. Ferreira, Courtenay T. Vaughan, Mark Swan
CLUSTER2
2008 Instrumentation and Analysis of MPI Queue Times on the SeaStar High-Performance Network
abstract
Understanding the communication behavior and network resource usage of parallel applications is critical to achieving high performance and scalability on systems with tens of thousands of network endpoints. The need for better understanding is not only driven by the desire to identify potential performance optimization opportunities for current networks, but is also a necessity for designing next-generation networking hardware. In this paper, we describe our approach to instrumenting the SeaStar interconnect on the Cray XT series of massively parallel processing machines to gather low-level network timing data. This data provides a new perspective on performance evaluation, both in terms of evaluating the resource usage patterns of applications as well as evaluating different implementation strategies in the network protocol stack.
Ron Brightwell, Kevin T. Pedretti, Kurt B. Ferreira
ICCCN2
2008 SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor
abstract
This paper describes SMARTMAP, an operating system technique that implements fixed offset virtual memory addressing. SMARTMAP allows the application processes on a multi-core processor to directly access each other's memory without the overhead of kernel involvement. When used to implement MPI, SMARTMAP eliminates all extraneous memory-to-memory copies imposed by UNIX-based shared memory strategies. In addition, SMARTMAP can easily support operations that UNIX-based shared memory cannot, such as direct, in-place MPI reduction operations and one-sided get/put operations. We have implemented SMARTMAP in the Catamount lightweight kernel for the Cray XT and modified MPI and Cray SHMEM libraries to use it. Micro-benchmark performance results show that SMARTMAP allows for significant improvements in latency, bandwidth, and small message rate on a quad-core processor.
Ron Brightwell, Kevin T. Pedretti, Trammell Hudson
SC2
2005 Implementation and Performance of Portals 3.3 on the Cray XT3
abstract
The Portals data movement interface was developed at Sandia National Laboratories in collaboration with the University of New Mexico over the last ten years. Portals is intended to provide the functionality necessary to scale a distributed memory parallel computing system to thousands of nodes. Previous versions of Portals ran on several large-scale machines, including a 1024-node nCUBE-2, a 1800-node Intel Paragon, and the 4500-node Intel ASCI Red machine. The latest version of Portals was initially developed for an 1800-node Linux/Myrinet cluster and has since been adopted by Cray as the lowest-level network programming interface for their XT3 platform. In this paper, we describe the implementation of Portals 3.3 on the Cray XT3 and present some initial performance results from several micro-benchmark tests. Despite some limitations, the implementation of Portals is able to achieve a zero-length one-way latency of under six microseconds and a uni-directional bandwidth of more than 1.1 GB/s
Ron Brightwell, Trammell Hudson, Kevin T. Pedretti, Rolf Riesen, Keith D. Underwood
CLUSTER3
2005 Gene transcript clustering: a comparison of parallel approaches
Todd E. Scheetz, Nishank Trivedi, Kevin T. Pedretti, Terry A. Braun, Thomas L. Casavant
Future Gener. Comput. Syst.3
2002 Cplant? Runtime System Support for Multi-Processor and Heterogeneous Compute Nodes
abstract
In this paper, we describe additions and modifications to the Computational Plant (Cplant/sup /spl trade//) system software to support multi-processor compute nodes and to support heterogeneous node types. We describe how these capabilities have been incorporated into our scalable runtime system and how these changes affect the interface seen by end users and application developers. We also discuss several important operating system and networking issues that can directly impact application performance. We present some initial performance metrics that indicate how our current implementation scales when multiple processes are running on a single node.
Kevin T. Pedretti, Ron Brightwell
CLUSTER1
2002 Parallel creation of non-redundant gene indices from partial mRNA transcripts
Nishank Trivedi, Jared Bischof, Kevin T. Pedretti, Todd E. Scheetz, Terry A. Braun, Chad A. Roberts, Natalie L. Robinson, Val C. Sheffield, Marcelo Bento Soares, Thomas L. Casavant
Future Gener. Comput. Syst.4
2001 Parallelization of local BLAST service on workstation clusters
R. C. Braun, Kevin T. Pedretti, Thomas L. Casavant, Todd E. Scheetz, Clayton L. Birkett, Chad A. Roberts
Future Gener. Comput. Syst.2