VLDB 2026 Research / reviewers in the wild / expert
Kevin T. Pedretti
dblp:91/1284 · also Kevin Pedretti
· DBLP profile ↗
28ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0002-1261-0178ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 2 first-author · 2 since 2021Computer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
High-performance computing · 32% Parallel and multicore computing · 18% Hardware reliability and fault tolerance · 16% | |
| Software engineering, system software, and programming languages
3 papers |
Operating systems · 74% Concurrent programming · 14% Runtime systems and virtual machines · 11% |
Topics — the 16 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › supercomputing
supercomputer deployment |
0.4 | 1 | 2020 | Chronicles of astra: challenges and lessons from the first petascale arm supercomputer · SC 2020 |
Interconnection networks and networks-on-chip
network congestion |
0.4 | 1 | 2019 | Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019 |
Parallel and multicore computing
task allocation |
0.4 | 1 | 2019 | Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019 |
Operating systems › kernel
lightweight kernel |
0.2 | 1 | 2015 | Achieving Performance Isolation with Lightweight Co-Kernels · HPDC 2015 |
Cloud and datacenter computing
performance isolation |
0.2 | 1 | 2015 | Achieving Performance Isolation with Lightweight Co-Kernels · HPDC 2015 |
Processor architecture and microarchitecture › instruction set architecture › RISC
ARM architecture |
0.2 | 1 | 2022 | Understanding Memory Failures on a Petascale Arm System · HPDC 2022 |
High-performance computing › supercomputing
exascale systems |
0.1 | 1 | 2011 | Evaluating the viability of process replication reliability for exascale systems · SC 2011 |
Distributed systems
fault tolerance |
0.1 | 1 | 2011 | Evaluating the viability of process replication reliability for exascale systems · SC 2011 |
Distributed systems
replication |
0.1 | 1 | 2011 | Evaluating the viability of process replication reliability for exascale systems · SC 2011 |
Distributed systems › replication
state machine replication |
0.1 | 1 | 2011 | Evaluating the viability of process replication reliability for exascale systems · SC 2011 |
Parallel and multicore computing › parallel programming models › message passing
MPI applications |
0.1 | 1 | 2019 | Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus Networks · IEEE Trans. Parallel Distributed Syst. 2019 |
Operating systems › resource management
memory management |
0.1 | 1 | 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008 |
Concurrent programming
shared memory |
0.1 | 1 | 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008 |
Parallel and multicore computing
MPI |
0.1 | 1 | 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008 |
Parallel and multicore computing
parallel programming models |
0.1 | 1 | 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008 |
Memory systems › memory management › virtual memory
virtual memory addressing |
0.1 | 1 | 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processor · SC 2008 |
Methods — techniques the papers use, named apart from their topics
containerization · 0.9statistical characterization · 0.6field failure data analysis · 0.6virtualization · 0.4co-kernel architecture · 0.4graph partitioning · 0.4geometric partitioning · 0.4simulation · 0.1modeling · 0.1empirical analysis · 0.1lightweight kernel · 0.1fixed offset virtual memory · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Characterizing the Impact of Job Execution on the Occurrence of Memory Failures on a Petascale HPC SystemabstractCharacterizing the reliability of current and recent high performance (HPC) systems is critical for forecasting how future systems may behave and informing the design of fault tolerance mechanisms. Although research has been conducted to understand memory failures, there are few examples where the occurrence of memory failures is considered in the broader context of system power, temperature, and the execution of user jobs. In this paper, we combine job data with existing power, temperature, and memory failure data collected on a petascale HPC system. By focusing on periods when jobs were running on the system, we identified trends that were not evident in earlier studies of this same data. The inclusion of job data also demonstrated how user behavior can affect the occurrence of memory failures. In conjunction with this paper, we have publicly released the job data used in this paper to complement existing publicly-available data from Astra regarding power, temperature, and the occurrence of correctable memory failures. Scott Levy, Joshua Hemmert, Kurt B. Ferreira, Kevin T. Pedretti |
SBAC-PAD | 4 |
| 2023 | Enabling power measurement and control on Astra: The first petascale Arm supercomputerabstractSummary Astra, deployed in 2018, was the first petascale supercomputer to utilize processors based on the ARM instruction set. The system was also the first under Sandia's Vanguard program which seeks to provide an evaluation vehicle for novel technologies that with refinement could be utilized in demanding, large‐scale HPC environments. In addition to ARM, several other important first‐of‐a‐kind developments were used in the machine, including new approaches to cooling the datacenter and machine. This article documents our experiences building a power measurement and control infrastructure for Astra. While this is often beyond the control of users today, the accurate measurement, cataloging, and evaluation of power, as our experiences show, is critical to the successful deployment of a large‐scale platform. While such systems exist in part for other architectures, Astra required new development to support the novel Marvell ThunderX2 processor used in compute nodes. In addition to documenting the measurement of power during system bring up and for subsequent on‐going routine use, we present results associated with controlling the power usage of the processor, an area which is becoming of progressively greater interest as data centers and supercomputing sites look to improve compute/energy efficiency and find additional sources for full system optimization. Ryan E. Grant, Simon D. Hammond, James H. Laros III, Michael J. Levenhagen, Stephen Olivier, Kevin T. Pedretti, Lee Ward, Andrew J. Younge |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Understanding Memory Failures on a Petascale Arm SystemabstractNew and novel HPC platforms provide interesting challenges and opportunities. Analysis of these systems can provide a better understanding of both the specific platform being studied as well as large-scale systems in general. Arm is one such architecture that has been explored in HPC for several years, however little is still known about its viability for supporting large-scale production workloads in terms of system reliability. The Astra system at Sandia National Laboratories was the first public peta-FLOPS Arm-based system on the Top500 and has been successfully running production HPC applications for a couple of years. In this paper, we analyze memory failure data collected from Astra while the system was in production running unclassified applications. This analysis revealed several interesting contributions related to both the Arm platform and to HPC systems in general. First, we outline the number of components replaced due to reliability issues in standing-up this first-of-its-kind, large-scale HPC system. We show the distribution differences between correctable DRAM faults and errors on Astra, showing that, not properly accounting for faults can lead to erroneous conclusions. Additionally, we characterize DRAM faults on the system and show contrary to existing work that memory faults are uniformly distributed across CPU socket, DRAM column, bank and rack region, but are not uniform across node, DIMM rank, DIMM slot on the motherboard, and system rack: some racks, ranks and DIMM slots experience more faults than others. Similarly, we show the impact of temperature and power on DRAM correctable errors. Finally, we make a detailed comparison of results presented here with the positional affects found in several previous large-scale reliability studies. The results of this analysis provide valuable guidance to organizations standing-up first-in- class platforms in HPC, organizations using Arm in HPC, and the entire large-scale HPC community in general. Kurt B. Ferreira, Scott Levy, Joshua Hemmert, Kevin T. Pedretti |
HPDC | 4 |
| 2020 | Chronicles of astra: challenges and lessons from the first petascale arm supercomputerabstractArm processors have been explored in HPC for several years, however there has not yet been a demonstration of viability for supporting large-scale production workloads. In this paper, we offer a retrospective on the process of bringing up Astra, the first Petascale supercomputer based on 64-bit Arm processors, and validating its ability to run production HPC applications. Through this process several immature technology gaps were addressed, including software stack enablement, Linux bugs at scale, thermal management issues, power management capabilities, and advanced container support. From this experience, several lessons learned are formulated that contributed to the successful deployment of Astra. These insights can be helpful to accelerate deploying and maturing other first-seen HPC technologies. With Astra now supporting many users running a diverse set of production applications at multi-thousand node scales, we believe this constitutes strong supporting evidence that Arm is a viable technology for even the largest-scale supercomputer deployments. Kevin T. Pedretti, Andrew J. Younge, Simon D. Hammond, James H. Laros III, Matthew L. Curry, Michael J. Aguilar, Robert J. Hoekstra, Ron Brightwell |
SC | 1 |
| 2019 | Geometric Mapping of Tasks to Processors on Parallel Computers with Mesh or Torus NetworksabstractWe present a new method for reducing parallel applications' communication time by mapping their MPI tasks to processors in a way that lowers the distance messages travel and the amount of congestion in the network. Assuming geometric proximity among the tasks is a good approximation of their communication interdependence, we use a geometric partitioning algorithm to order both the tasks and the processors, assigning task parts to the corresponding processor parts. In this way, interdependent tasks are assigned to “nearby” cores in the network. We also present a number of algorithmic optimizations that exploit specific features of the network or application to further improve the quality of the mapping. We specifically address the case of sparse node allocation, where the nodes assigned to a job are not necessarily located in a contiguous block nor within close proximity to each other in the network. However, our methods generalize to contiguous allocations as well, and results are shown for both contiguous and non-contiguous allocations. We show that, for the structured finite difference mini-application MiniGhost, our mapping methods reduced communication time up to 75 percent relative to MiniGhost's default mapping on 128K cores of a Cray XK7 with sparse allocation. For the atmospheric modeling code E3SM/HOMME, our methods reduced communication time up to 31% on 16K cores of an IBM BlueGene/Q with contiguous allocation. Mehmet Deveci, Karen D. Devine, Kevin T. Pedretti, Mark A. Taylor, Sivasankaran Rajamanickam, Ümit V. Çatalyürek |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Large-Scale System Monitoring Experiences and RecommendationsabstractMonitoring of High Performance Computing (HPC) platforms is critical to successful operations, can provide insights into performance-impacting conditions, and can inform methodologies for improving science throughput. However, monitoring systems are not generally considered core capabilities in system requirements specifications nor in vendor development strategies. In this paper we present work performed at a number of large-scale HPC sites towards developing monitoring capabilities that fill current gaps in ease of problem identification and root cause discovery. We also present our collective views, based on the experiences presented, on needs and requirements for enabling development by vendors or users of effective sharable end-to-end monitoring capabilities. Ville Ahlgren, Stefan Andersson, Jim M. Brandt, Nicholas Cardo, Sudheer Chunduri, Jeremy Enos, Parks Fields, Ann C. Gentile, Richard A. Gerber, Michael Gienger, Joe Greenseid, Annette Greiner, Bilel Hadri, Dennis Hoppe, Urpo Kaila, Kaki Kelly, Mark Klein 0002, Alex Kristiansen, Stephen Leak, Mike Mason, Kevin T. Pedretti, Jean-Guillaume Piccinali, Jason Repik, Jim Rogers, Susanna Salminen, Michael T. Showerman, Cary Whitney, Jim Williams |
CLUSTER | 22 |
| 2018 | Characterizing MPI matching via trace-based simulation
Kurt B. Ferreira, Scott Levy, Kevin T. Pedretti, Ryan E. Grant |
Parallel Comput. | 3 |
| 2017 | A Tale of Two Systems: Using Containers to Deploy HPC Applications on Supercomputers and CloudsabstractContainerization, or OS-level virtualization has taken root within the computing industry. However, container utilization and its impact on performance and functionality within High Performance Computing (HPC) is still relatively undefined. This paper investigates the use of containers with advanced supercomputing and HPC system software. With this, we define a model for parallel MPI application DevOps and deployment using containers to enhance development effort and provide container portability from laptop to clouds or supercomputers. In this endeavor, we extend the use of Sin- gularity containers to a Cray XC-series supercomputer. We use the HPCG and IMB benchmarks to investigate potential points of overhead and scalability with containers on a Cray XC30 testbed system. Furthermore, we also deploy the same containers with Docker on Amazon's Elastic Compute Cloud (EC2), and compare against our Cray supercomputer testbed. Our results indicate that Singularity containers operate at native performance when dynamically linking Cray's MPI libraries on a Cray supercomputer testbed, and that while Amazon EC2 may be useful for initial DevOps and testing, scaling HPC applications better fits supercomputing resources like a Cray. Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Ron Brightwell |
CloudCom | 2 |
| 2017 | Enabling Diverse Software Stacks on Supercomputers Using High Performance Virtual ClustersabstractWhile large-scale simulations have been the hallmark of the High Performance Computing (HPC) community for decades, Large Scale Data Analytics (LSDA) workloads are gaining attention within the scientific community not only as a processing component to large HPC simulations, but also as standalone scientific tools for knowledge discovery. With the path towards Exascale, new HPC runtime systems are also emerging in a way that differs from classical distributed computing models. However, system software for such capabilities on the latest extreme-scale DOE supercomputing needs to be enhanced to more appropriately support these types of emerging software ecosystems. In this paper, we propose the use of Virtual Clusters on advanced supercomputing resources to enable systems to support not only HPC workloads, but also emerging big data stacks. Specifically, we have deployed the KVM hypervisor within Cray's Compute Node Linux on a XC-series supercomputer testbed. We also use libvirt and QEMU to manage and provision VMs directly on compute nodes, leveraging Ethernet-over-Aries network emulation. To our knowledge, this is the first known use of KVM on a true MPP supercomputer. We investigate the overhead our solution using HPC benchmarks, both evaluating single-node performance as well as weak scaling of a 32-node virtual cluster. Overall, we find single node performance of our solution using KVM on a Cray is very efficient with near-native performance. However overhead increases by up to 20% as virtual cluster size increases, due to limitations of the Ethernet-over-Aries bridged network. Furthermore, we deploy Apache Spark with large data analysis workloads in a Virtual Cluster, effectively demonstrating how diverse software ecosystems can be supported by High Performance Virtual Clusters. Andrew J. Younge, Kevin T. Pedretti, Ryan E. Grant, Brian L. Gaines, Ron Brightwell |
CLUSTER | 2 |
| 2016 | Sweet Spots and Limits for VirtualizationabstractThis year at VEE, we added a panel to discuss the state of virtualization: what problems are solved? what problems are important? and what problems may not be worth solving? The panelist are experts in areas ranging from hardware virtualization up to language-level virtualization. Carl A. Waldspurger, Emery D. Berger, Abhishek Bhattacharjee, Kevin T. Pedretti, Simon Peter 0001, Christopher J. Rossbach |
VEE | 4 |
| 2015 | Achieving Performance Isolation with Lightweight Co-KernelsabstractPerformance isolation is emerging as a requirement for High Performance Computing (HPC) applications, particularly as HPC architectures turn to in situ data processing and application composition techniques to increase system throughput. These approaches require the co-location of disparate workloads on the same compute node, each with different resource and runtime requirements. In this paper we claim that these workloads cannot be effectively managed by a single Operating System/Runtime (OS/R). Therefore, we present Pisces, a system software architecture that enables the co-existence of multiple independent and fully isolated OS/Rs, or enclaves, that can be customized to address the disparate requirements of next generation HPC workloads. Each enclave consists of a specialized lightweight OS co-kernel and runtime, which is capable of independently managing partitions of dynamically assigned hardware resources. Contrary to other co-kernel approaches, in this work we consider performance isolation to be a primary requirement and present a novel co-kernel architecture to achieve this goal. We further present a set of design requirements necessary to ensure performance isolation, including: (1) elimination of cross OS dependencies, (2) internalized management of I/O, (3) limiting cross enclave communication to explicit shared memory channels, and (4) using virtualization techniques to provide missing OS features. The implementation of the Pisces co-kernel architecture is based on the Kitten Lightweight Kernel and Palacios Virtual Machine Monitor, two system software architectures designed specifically for HPC systems. Finally we will show that lightweight isolated co-kernels can provide better performance for HPC applications, and that isolated virtual machines are even capable of outperforming native environments in the presence of competing workloads. Jiannan Ouyang, Brian Kocoloski, Jack Lange, Kevin T. Pedretti |
HPDC | 4 |
| 2014 | Demonstrating improved application performance using dynamic monitoring and task mappingabstractThis work demonstrates the integration of monitoring, analysis, and feedback to perform application-to-resource mapping that adapts to both static architecture features and dynamic resource state. In particular, we present a framework for mapping MPI tasks to compute resources based on run-time analysis of system-wide network data, architecture-specific routing algorithms, and application communication patterns. We address several challenges. Within each node, we collect local utilization data. We consolidate that information to form a global view of system performance, accounting for system-wide factors including competing applications. We provide an interface for applications to query the global information. Then we exploit the system information to change the mapping of tasks to nodes so that system bottlenecks are avoided. We demonstrate the benefit of this monitoring and feedback by remapping MPI tasks based on route-length, bandwidth, and credit-stalls metrics for a parallel sparse matrix-vector multiplication kernel. In the best case, remapping based on dynamic network information in a congested environment recovered 48.9% of the time lost to congestion, reducing matrix-vector multiplication time by 7.8%. Our experiments focus on the Cray XE/XK platform, but the integration concepts are generally applicable to any platform for which applicable metrics and route knowledge can be obtained. Jim M. Brandt, Karen D. Devine, Ann C. Gentile, Kevin T. Pedretti |
CLUSTER | 4 |
| 2014 | Exploiting Geometric Partitioning in Task Mapping for Parallel ComputersabstractWe present a new method for mapping applications' MPI tasks to cores of a parallel computer such that communication and execution time are reduced. We consider the case of sparse node allocation within a parallel machine, where the nodes assigned to a job are not necessarily located within a contiguous block nor within close proximity to each other in the network. The goal is to assign tasks to cores so that interdependent tasks are performed by "nearby" cores, thus lowering the distance messages must travel, the amount of congestion in the network, and the overall cost of communication. Our new method applies a geometric partitioning algorithm to both the tasks and the processors, and assigns task parts to the corresponding processor parts. We show that, for the structured finite difference mini-app Mini Ghost, our mapping method reduced execution time 34% on average on 65,536 cores of a Cray XE6. In a molecular dynamics mini-app, Mini MD, our mapping method reduced communication time by 26% on average on 6144 cores. We also compare our mapping with graph-based mappings from the LibTopoMap library and show that our mappings reduced the communication time on average by 15% in MiniGhost and 10% in MiniMD. Mehmet Deveci, Sivasankaran Rajamanickam, Vitus J. Leung, Kevin T. Pedretti, Stephen Olivier, David P. Bunde, Ümit V. Çatalyürek, Karen D. Devine |
IPDPS | 4 |
| 2014 | Exascale design space exploration and co-design
Sudip S. Dosanjh, Richard F. Barrett, Douglas Doerfler, Simon D. Hammond, Karl S. Hemmert, Michael A. Heroux, Paul T. Lin, Kevin T. Pedretti, Arun Rodrigues, Timothy G. Trucano, Justin Luitjens |
Future Gener. Comput. Syst. | 8 |
| 2012 | Application-driven analysis of two generations of capability computing: the transition to multicore processorsabstractSUMMARY Multicore processors form the basis of most traditional high performance parallel processing architectures. Early experiences with these computers showed significant performance problems, both with regard to computation and inter‐process communication. The transition from Purple, an IBM POWER5‐based machine, to Cielo, a Cray XE6, as the main capability computing platform for the United States Department of Energy's Advanced Simulation and Computing campaign provides an opportunity to reexamine these issues after experiences with a few generations of multicore‐based machines. Experiences with Purple identified some important characteristics that led to strong performance of complex scientific application programs at very large scales. Herein, we compare the performance of some Advanced Simulation and Computing mission critical applications at capability scale across this transition to multicore processors. Copyright © 2012 John Wiley & Sons, Ltd. Mahesh Rajan, Courtenay T. Vaughan, Douglas Doerfler, Richard F. Barrett, Paul T. Lin, Kevin T. Pedretti, Karl S. Hemmert |
Concurr. Comput. Pract. Exp. | 6 |
| 2011 | The Impact of Injection Bandwidth Performance on Application Scalability
Kevin T. Pedretti, Ron Brightwell, Douglas Doerfler, Karl S. Hemmert, James H. Laros III |
EuroMPI | 1 |
| 2011 | Evaluating the viability of process replication reliability for exascale systemsabstractAs high-end computing machines continue to grow in size, issues such as fault tolerance and reliability limit application scalability. Current techniques to ensure progress across faults, like checkpoint-restart, are increasingly problematic at these scales due to excessive overheads predicted to more than double an application's time to solution. Replicated computing techniques, particularly state machine replication, long used in distributed and mission critical systems, have been suggested as an alternative to checkpoint-restart. In this paper, we evaluate the viability of using state machine replication as the primary fault tolerance mechanism for upcoming exascale systems. We use a combination of modeling, empirical analysis, and simulation to study the costs and benefits of this approach in comparison to checkpoint/restart on a wide range of system parameters. These results, which cover different failure distributions, hardware mean time to failures, and I/O bandwidths, show that state machine replication is a potentially useful technique for meeting the fault tolerance demands of HPC applications on future exascale platforms. Kurt B. Ferreira, Jon Stearley, James H. Laros III, Ron A. Oldfield, Kevin T. Pedretti, Ron Brightwell, Rolf Riesen, Patrick G. Bridges, Dorian C. Arnold |
SC | 5 |
| 2011 | Minimal-overhead virtualization of a large scale supercomputerabstractVirtualization has the potential to dramatically increase the usability and reliability of high performance computing (HPC) systems. However, this potential will remain unrealized unless overheads can be minimized. This is particularly challenging on large scale machines that run carefully crafted HPC OSes supporting tightly-coupled, parallel applications. In this paper, we show how careful use of hardware and VMM features enables the virtualization of a large-scale HPC system, specifically a Cray XT4 machine, with < = 5% overhead on key HPC applications, microbenchmarks, and guests at scales of up to 4096 nodes. We describe three techniques essential for achieving such low overhead: passthrough I/O, workload-sensitive selection of paging mechanisms, and carefully controlled preemption. These techniques are forms of symbiotic virtualization, an approach on which we elaborate. Jack Lange, Kevin T. Pedretti, Peter A. Dinda, Patrick G. Bridges, Chang Bae, Philip Soltero, Alex Merritt |
VEE | 2 |
| 2010 | The Impact of System Design Parameters on Application Noise SensitivityabstractOperating system noise, or “jitter,” is a key limiter of application scalability in high end computing systems. Several studies have attempted to quantify the sources and effects of system interference, though few of these studies show the influence that architectural and system characteristics have on the impact of OS noise at scale. In this paper, we examine the impact of three such system properties: platform balance, “noisy” node distribution, and non-blocking collective operations. Using a previouslydeveloped noise injection tool, we explore how the impact of noise varies with these platform characteristics. We provide detailed performance results that indicate that a system with relatively less network bandwidth is able to absorb more noise than a system with more network bandwidth. Our results also show that application performance can be significantly degraded by only a subset of noisy nodes. Furthermore, the placement of the noisy nodes is also important, especially for applications that make substantial use of collective communication operations that are tree-based. Lastly, performance results indicate that nonblocking collective operations have the ability to greatly mitigate the impact of OS interference. Combined, these results show that the impact of OS noise is not solely a property of application communication behavior, but is also influenced by other properties of the system architecture and system software environment. Kurt B. Ferreira, Patrick G. Bridges, Ron Brightwell, Kevin T. Pedretti |
CLUSTER | 4 |
| 2010 | Palacios and Kitten: New high performance operating systems for scalable virtualized and native supercomputingabstractPalacios is a new open-source VMM under development at Northwestern University and the University of New Mexico that enables applications executing in a virtualized environment to achieve scalable high performance on large machines. Palacios functions as a modularized extension to Kitten, a high performance operating system being developed at Sandia National Laboratories to support large-scale supercomputing applications. Together, Palacios and Kitten provide a thin layer over the hardware to support full-featured virtualized environments alongside Kitten's lightweight native environment. Palacios supports existing, unmodified applications and operating systems by using the hardware virtualization technologies in recent AMD and Intel processors. Additionally, Palacios leverages Kitten's simple memory management scheme to enable low-overhead pass-through of native devices to a virtualized environment. We describe the design, implementation, and integration of Palacios and Kitten. Our benchmarks show that Palacios provides near native (within 5%), scalable performance for virtualized environments running important parallel applications. This new architecture provides an incremental path for applications to use supercomputers, running specialized lightweight host operating systems, that is not significantly performance-compromised. Jack Lange, Kevin T. Pedretti, Trammell Hudson, Peter A. Dinda, Lei Xia 0001, Patrick G. Bridges, Andy Gocke, Steven Jaconette, Michael J. Levenhagen, Ron Brightwell |
IPDPS | 2 |
| 2009 | Topics on measuring real power usage on high performance computing platformsabstractPower has recently been recognized as one of the major obstacles in fielding a Peta-FLOPs class system. To reach Exa-FLOPs, the challenge will certainly be compounded. In this paper we will discuss a number of High Performance Computing power related topics. We first describe our implementation of a scalable power measurement framework that has enabled us to examine real power use (current draw). [Using this framework, samples were obtained at a per-node (socket) granularity, at frequencies of up to 100 samples per second.] Additionally, we describe how we applied this capability to implement power conserving measures on our Catamount Light Weight Kernel, where we achieved an 80% improvement. This ability has enabled us to quantify the amount of energy used by applications and to contrast application energy use between a Light Weight and General Purpose operating system. Finally, we show application energy use increases proportionally with the increase in run-time due to operating system noise. Areas of future interest will also be discussed. James H. Laros III, Kevin T. Pedretti, Suzanne M. Kelly, John P. Vandyke, Kurt B. Ferreira, Courtenay T. Vaughan, Mark Swan |
CLUSTER | 2 |
| 2008 | Instrumentation and Analysis of MPI Queue Times on the SeaStar High-Performance NetworkabstractUnderstanding the communication behavior and network resource usage of parallel applications is critical to achieving high performance and scalability on systems with tens of thousands of network endpoints. The need for better understanding is not only driven by the desire to identify potential performance optimization opportunities for current networks, but is also a necessity for designing next-generation networking hardware. In this paper, we describe our approach to instrumenting the SeaStar interconnect on the Cray XT series of massively parallel processing machines to gather low-level network timing data. This data provides a new perspective on performance evaluation, both in terms of evaluating the resource usage patterns of applications as well as evaluating different implementation strategies in the network protocol stack. Ron Brightwell, Kevin T. Pedretti, Kurt B. Ferreira |
ICCCN | 2 |
| 2008 | SMARTMAP: operating system support for efficient data sharing among processes on a multi-core processorabstractThis paper describes SMARTMAP, an operating system technique that implements fixed offset virtual memory addressing. SMARTMAP allows the application processes on a multi-core processor to directly access each other's memory without the overhead of kernel involvement. When used to implement MPI, SMARTMAP eliminates all extraneous memory-to-memory copies imposed by UNIX-based shared memory strategies. In addition, SMARTMAP can easily support operations that UNIX-based shared memory cannot, such as direct, in-place MPI reduction operations and one-sided get/put operations. We have implemented SMARTMAP in the Catamount lightweight kernel for the Cray XT and modified MPI and Cray SHMEM libraries to use it. Micro-benchmark performance results show that SMARTMAP allows for significant improvements in latency, bandwidth, and small message rate on a quad-core processor. Ron Brightwell, Kevin T. Pedretti, Trammell Hudson |
SC | 2 |
| 2005 | Implementation and Performance of Portals 3.3 on the Cray XT3abstractThe Portals data movement interface was developed at Sandia National Laboratories in collaboration with the University of New Mexico over the last ten years. Portals is intended to provide the functionality necessary to scale a distributed memory parallel computing system to thousands of nodes. Previous versions of Portals ran on several large-scale machines, including a 1024-node nCUBE-2, a 1800-node Intel Paragon, and the 4500-node Intel ASCI Red machine. The latest version of Portals was initially developed for an 1800-node Linux/Myrinet cluster and has since been adopted by Cray as the lowest-level network programming interface for their XT3 platform. In this paper, we describe the implementation of Portals 3.3 on the Cray XT3 and present some initial performance results from several micro-benchmark tests. Despite some limitations, the implementation of Portals is able to achieve a zero-length one-way latency of under six microseconds and a uni-directional bandwidth of more than 1.1 GB/s Ron Brightwell, Trammell Hudson, Kevin T. Pedretti, Rolf Riesen, Keith D. Underwood |
CLUSTER | 3 |
| 2005 | Gene transcript clustering: a comparison of parallel approaches
Todd E. Scheetz, Nishank Trivedi, Kevin T. Pedretti, Terry A. Braun, Thomas L. Casavant |
Future Gener. Comput. Syst. | 3 |
| 2002 | Cplant? Runtime System Support for Multi-Processor and Heterogeneous Compute NodesabstractIn this paper, we describe additions and modifications to the Computational Plant (Cplant/sup /spl trade//) system software to support multi-processor compute nodes and to support heterogeneous node types. We describe how these capabilities have been incorporated into our scalable runtime system and how these changes affect the interface seen by end users and application developers. We also discuss several important operating system and networking issues that can directly impact application performance. We present some initial performance metrics that indicate how our current implementation scales when multiple processes are running on a single node. Kevin T. Pedretti, Ron Brightwell |
CLUSTER | 1 |
| 2002 | Parallel creation of non-redundant gene indices from partial mRNA transcripts
Nishank Trivedi, Jared Bischof, Kevin T. Pedretti, Todd E. Scheetz, Terry A. Braun, Chad A. Roberts, Natalie L. Robinson, Val C. Sheffield, Marcelo Bento Soares, Thomas L. Casavant |
Future Gener. Comput. Syst. | 4 |
| 2001 | Parallelization of local BLAST service on workstation clusters
R. C. Braun, Kevin T. Pedretti, Thomas L. Casavant, Todd E. Scheetz, Clayton L. Birkett, Chad A. Roberts |
Future Gener. Comput. Syst. | 2 |