Jack Lange

dblp:65/3740 · also John Jack Lange, John Lange, John R. Lange · DBLP profile ↗
← Back
29ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0003-0616-7437ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 6 first-author · 4 since 2021Security and privacy · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 An Evaluation of the Effect of Network Cost Optimization for Leadership Class Supercomputers
abstract
Dragonfly-based networks are an extensively deployed network topology in large-scale high-performance computing due to their cost-effectiveness and efficiency. The US will soon have three Exascale supercomputers for leadership class workloads deployed using dragonfly networks. Compared to indirect networks of similar scale, the dragonfly network has considerably reduced cable lengths, cable counts, and switch counts, resulting in significant network cost savings for a given system size, however, these cost reductions result in reduced global minimal paths and more challenging routing. Additionally, large scale dragonfly networks often require a taper at the global link level, resulting in less bisection bandwidth than is achievable in other traditional non-blocking topologies of equivalent scale. While dragonfly networks have been extensively studied, they have yet to be fully evaluated in an extreme scale (i.e., exascale) system that targets capability workloads. In this paper, we present the results of the first large scale evaluation of a dragonfly network on an exascale system (Frontier) and compare its behavior to a similar scale fat-tree network on a previous generation TOP500 system (Summit). This evaluation aims to determine the effect of network cost optimizations by measuring a tapered topology’s impact on capability workloads. Our evaluation is based on a collection of synthetic microbenchmarks, mini-apps, and full scale applications. It compares the scaling efficiencies of each benchmark between the dragonfly-based Frontier and the fat-tree-based Summit systems. Our results show that a dragonfly network is $\sim \mathbf{3 0 \%}$ more cost efficient than a fat-tree topology, which amortizes to $\sim 3 \%$ of an exascale system cost. Furthermore, while tapered dragonfly networks impose significant tradeoffs, the impacts are not as broad as initially thought and are mostly seen in applications with global communication patterns, particularly all-to-all (e.g., FFT-based algorithms), but also local communication patterns (e.g., nearest-neighbor algorithms) that are sensitive to network performance variability.
Awais Khan 0002, Jack Lange, Nick Hagerty, Edwin F. Posada, John K. Holmen, James B. White, James Austin Harris, Verónica G. Vergara Larrea, Christopher Zimmer 0001, Scott Atchley
SC2
2023 Frontier: Exploring Exascale
abstract
As the US Department of Energy (DOE) computing facilities began deploying petascale systems in 2008, DOE was already setting its sights on exascale. In that year, DARPA published a report on the feasibility of reaching exascale. The report authors identified several key challenges in the pursuit of exascale including power, memory, concurrency, and resiliency. That report informed the DOE's computing strategy for reaching exascale. With the deployment of Oak Ridge National Laboratory's Frontier supercomputer, we have officially entered the exascale era. In this paper, we discuss Frontier's architecture, how it addresses those challenges, and describe some early application results from Oak Ridge Leadership Computing Facility's Center of Excellence and the Exascale Computing Project.
Scott Atchley, Christopher Zimmer 0001, Jack Lange, David E. Bernholdt, Verónica G. Vergara Larrea, Michael J. Brim, Reuben D. Budiardja, Sunita Chandrasekaran, Markus Eisenbach 0002, Thomas M. Evans 0001, Matthew Ezell, Nicholas Frontiere, Antigoni Georgiadou, Joseph Glenski, Philipp Grete, Steven P. Hamilton, John K. Holmen, Axel Huebl, Daniel A. Jacobson, Wayne Joubert, Kim H. McMahon, Elia Merzari, Stan G. Moore, Andrew Myers 0001, Stephen Nichols, Sarp Oral, Thomas Papatheodore, Danny Perez, David M. Rogers 0001, Evan Schneider, Jean-Luc Vay, P. K. Yeung
SC3
2022 Lifting and Dropping VMs to Dynamically Transition Between Time- and Space-sharing for Large-Scale HPC Systems
abstract
As HPC environments increasingly integrate with edge based systems, system architectures will need to handle a broader class of workloads and scheduling requirements. One result of this shift will be the need to simultaneously support bulk-synchronous parallel (BSP) and on-demand service based applications on the same infrastructure. This in turn will require that future resource management approaches utilize both space-shared as well as time-shared resource scheduling strategies. In this work we introduce the concept of "VM-lifting'' (and its inverse "VM-Dropping'') which allows dynamically switching an HPC workload between space-shared and time-shared scheduling regimes. Our work targets co-kernel based HPC system software environments, in which multiple specialized OS kernels execute natively on dedicated physical resource partitions inside a single compute node. With VM-lifting, a native co-kernel can be migrated at runtime to and from locally hosted Virtual Machine Environments due to changing scheduling requirements of the node. This allows an HPC node to be dynamically (re-)configured as either a time-shared Infrastructure-as-a-Service (IaaS) resource or a dedicated space shared resource based on the current workload demands. We have implemented this approach in the context of the Hobbes Exascale System Software stack and have demonstrated that a node can be reconfigured with minimal impact on the running applications.
Nicholas Gordon, Jack Lange
HPDC2
2022 Deadlock prediction via generalized dependency
abstract
Deadlocks are notorious bugs in multithreaded programs, causing serious reliability issues. However, they are difficult to be fully expunged before deployment, as their appearances typically depend on specific inputs and thread schedules, which require the assistance of dynamic tools. However, existing deadlock detection tools mainly focus on locks, but cannot detect deadlocks related to condition variables. This paper presents a novel approach to fill this gap. It extends the classic lock dependency to generalized dependency by abstracting the signal for the condition variable as a special resource so that communication deadlocks can be modeled as hold-and-wait cycles as well. It further designs multiple practical mechanisms to record and analyze generalized dependencies. In the end, this paper presents the implementation of the tool, called UnHang. Experimental results on real applications show that UnHang is able to find all known deadlocks and uncover two new deadlocks. Overall, UnHang only imposes around 3% performance overhead and 8% memory overhead, making it a practical tool for the deployment environment.
Jinpeng Zhou, Hanmei Yang, Jack Lange, Tongping Liu
ISSTA3
2021 Covirt: Lightweight Fault Isolation and Resource Protection for Co-Kernels
abstract
The challenges of the exascale era have generated a number of advancements in HPC systems software, with co-kernel architectures emerging as one such novel approach for HPC operating system and runtime (OS/R) design. Cokernels function by running multiple specialized, lightweight OS kernels natively on the same host as a general purpose OS/R. These specialized kernels are able to provide optimized OS/R environments for HPC applications while still retaining access to the full feature set of the co-running general purpose OS/R. While co-kernels are able to effectively optimize for performance, they generally lack effective mechanisms for cross OS/R fault isolation and resource protection. In this paper we present Covirt, a lightweight OS/R protection layer that leverages the hardware virtualization features found on modern CPUs. Covirt interposes a minimal hypervisor layer between a co-kernel OS/R and hardware to prevent OS level faults from impacting other OS/Rs running on the same system. Covirt is different from other virtualization-based approaches due to the level of integration necessary between the co-kernel instances, requiring the support of higher level semantic interfaces between the different OS/Rs. Covirt features a split architecture consisting of a hypervisor and controller module that continuously monitors changes to the underlying resource partitioning and translates those events to hypervisor configuration changes. We have implemented a prototype of Covirt in the context of the Hobbes exascale OS/R stack, specifically targeting the Pisces co-kernel framework and Kitten Lightweight Kernel. Our evaluation shows that Covirt is able to add fault isolation for memory and interrupt processing with minimal performance overheads.
Nicholas Gordon, Jack Lange
IPDPS2
2019 NeXUS: Practical and Secure Access Control on Untrusted Storage Platforms using Client-Side SGX
abstract
With the rising popularity of file-sharing services such as Google Drive and Dropbox in the workflows of individuals and corporations alike, the protection of client-outsourced data from unauthorized access or tampering remains a major security concern. Existing cryptographic solutions to this problem typically require server-side support, involve non-trivial key management on the part of users, and suffer from severe re-encryption penalties upon access revocations. This combination of performance overheads and management burdens makes this class of solutions undesirable in situations where performant, platform-agnostic, dynamic sharing of user content is required. We present NEXUS, a stackable filesystem that leverages trusted hardware to provide confidentiality and integrity for user files stored on untrusted platforms. NEXUS is explicitly designed to balance security, portability, and performance: it supports dynamic sharing of protected volumes on any platform exposing a file access API without requiring server-side support, enables the use of fine-grained access control policies to allow for selective sharing, and avoids the key revocation and file re-encryption overheads associated with other cryptographic approaches to access control. This combination of features is made possible by the use of a client-side Intel SGX enclave that is used to protect and share NEXUS volumes, ensuring that cryptographic keys never leave enclave memory and obviating the need to reencrypt files upon revocation of access rights. We implemented a NEXUS prototype that runs on top of the AFS filesystem and show that it incurs ×2 overhead for a variety of common file and database operations.
Judicael Briand Djoko, Jack Lange, Adam J. Lee
DSN2
2018 Varbench: an Experimental Framework to Measure and Characterize Performance Variability
abstract
Performance variability is a major problem for extreme scale parallel computing applications that rely on bulk synchronization and collective communication. While this problem is most prominent in the context of exascale systems, it is increasingly impacting other communities such as machine learning and graph analytics. In this paper, we present an experimental performance analysis framework called varbench that is designed to precisely measure the prevalence of performance variability in a system, as well as to support workload characterization with respect to how and when a workload generates variability. We demonstrate several of varbench's capabilities as they pertain to exascale-class systems, including its utility for discovering architectural trends, for performing cross-architectural comparisons, and for understanding key statistical properties of performance distributions that have implications for how system software should be designed to mitigate variability.
Brian Kocoloski, Jack Lange
ICPP2
2017 Harvesting Underutilized Resources to Improve Responsiveness and Tolerance to Crash and Silent Faults for Data-Intensive Applications
abstract
Low latency is critical for emerging data-intensive real-time analytic and interactive single-wave applications. As the complexity and heterogeneity of cloud computing continue to increase, so will the frequency of errors, which will manifest themselves in unpredictable ways with a significant impact on the correctness of the computing and responsiveness of cloud services. In this paper, we propose a new data-centric computational model to improve the responsiveness of the data-intensive applications to crash faults and augment its ability to deal with silent errors to ensure computational accuracy. The basic tenet of the proposed model is a task replication scheme, which interweaves the processing of a replicated data split among multiple distributed tasks, with each task consuming data at a different offset. In the absence of a failure, the concurrent execution of the tasks ensures complete processing of the data split, with a significant reduction in the total execution time. In the case of an error, however, the remaining tasks take over the execution of the unfinished work and finish on time. The proposed scheme also guarantees timely detection and correction of silent data corruptions along with crash faults. We demonstrate the effectiveness of our scheme by extending Hadoop's MapReduce code base as a case study. Results show a performance improvement of 50% over Hadoop's Speculative Execution when dealing with crash-faults and an improvement of 33% when dealing with silent errors in case of no failure.
Debashis Ganguly, Mohammad H. Mofrad, Taieb Znati, Rami G. Melhem, Jack Lange
CLOUD5
2016 Argo: Architecture-aware graph partitioning
abstract
The increasing popularity and ubiquity of various large graph datasets has caused renewed interest for graph partitioning. Existing graph partitioners either scale poorly against large graphs or disregard the impact of the underlying hardware topology. A few solutions have shown that the nonuniform network communication costs may affect the performance greatly. However, none of them considers the impact of resource contention on the memory subsystems (e.g., LLC and Memory Controller) of modern multicore clusters. They all neglect the fact that the bandwidth of modern high-speed networks (e.g., Infiniband) has become comparable to that of the memory subsystems. In this paper, we provide an in-depth analysis, both theoretically and experimentally, on the contention issue for distributed workloads. We found that the slowdown caused by the contention can be as high as 11x. We then design an architecture-aware graph partitioner, Argo, to allow the full use of all cores of multicore machines without suffering from either the contention or the communication heterogeneity issue. Our experimental study showed (1) the effectiveness of Argo, achieving up to 12x speedups on three classic workloads: Breadth First Search, Single Source Shortest Path, and PageRank; and (2) the scalability of Argo in terms of both graph size and the number of partitions on two billion-edge real-world graphs.
Angen Zheng, Alexandros Labrinidis, Panos K. Chrysanthis, Jack Lange
IEEE BigData4
2016 A Case for Criticality Models in Exascale Systems
abstract
Performance variation is a significant problem for large scale HPC systems and will increase on future exascale systems. In this work, we show that performance variation impacts the performance and energy efficiency of contemporary large-scale computing systems in highly temporally inconsistent ways. We thus present a case for criticality models, a learning based mechanism that allows a system to generate holistic models of performance variation as it occurs during application runtime. Criticality models are designed to provide a mechanism by which applications can detect performance variation at runtime and take action to mitigate its effects. We present a promising preliminary analysis of criticality models on a small scale cluster. Our results demonstrate that models based on logistic regression scan accurately model criticality at this scale.
Brian Kocoloski, Leonardo Piga, Wei Huang 0004, Indrani Paul, Jack Lange
CLUSTER5
2016 Shoot4U: Using VMM Assists to Optimize TLB Operations on Preempted vCPUs
abstract
Virtual Machine based approaches to workload consolidation, as seen in IaaS cloud as well as datacenter platforms, have long had to contend with performance degradation caused by synchronization primitives inside the guest environments. These primitives can be affected by virtual CPU preemptions by the host scheduler that can introduce delays that are orders of magnitude longer than those primitives were designed for. While a significant amount of work has focused on the behavior of spinlock primitives as a source of these performance issues, spinlocks do not represent the entirety of synchronization mechanisms that are susceptible to scheduling issues when running in a virtualized environment. In this paper we address the virtualized performance issues introduced by TLB shootdown operations. Our profiling study, based on the PARSEC benchmark suite, has shown that up to 64% of a VM's CPU time can be spent on TLB shootdown operations under certain workloads. In order to address this problem, we present a paravirtual TLB shootdown scheme named Shoot4U. Shoot4U completely eliminates TLB shootdown preemptions by invalidating guest TLB entries from the VMM and allowing guest TLB shootdown operations to complete without waiting for remote virtual CPUs to be scheduled. Our performance evaluation using the PARSEC benchmark suite demonstrates that Shoot4U can reduce benchmark runtime by up to 85% compared an unmodified Linux kernel, and up to 44% over a state-of-the-art paravirtual TLB shootdown scheme.
Jiannan Ouyang, Jack Lange, Haoqiang Zheng
VEE2
2016 Lightweight Memory Management for High Performance Applications in Consolidated Environments
abstract
Linux-based operating systems and runtimes (OS/Rs) have emerged as the environments of choice for the majority of HPC systems. While Linux-based OS/Rs have advantages such as extensive feature sets and developer familiarity, these features come at the cost of additional system overhead. In contrast to Linux, there is a substantial history of work in the HPC community focused on lightweight OS/Rs that provide scalable and consistent performance for HPC applications, but lack many of the features offered by commodity OS/Rs. In this paper, we propose to bridge the gap between LWKs and commodity OS/Rs by selectively providing a lightweight memory subsystem for HPC applications in a commodity OS/R where concurrently executing a diverse range of workloads is commonplace. Our system HPMMAP provides lightweight memory performance transparently to HPC applications by bypassing Linux's memory management layer. Using HPMMAP, HPC applications achieve consistent performance while the same local compute nodes execute competing workloads likely to be found in HPC clusters and “in-situ” workload deployments. Our approach is dynamically configurable at runtime, and requires no resources when not in use. We show that HPMMAP can decrease variance and reduce application runtime by up to 50 percent when executing a co-located competing commodity workload.
Brian Kocoloski, Jack Lange
IEEE Trans. Parallel Distributed Syst.2
2015 XEMEM: Efficient Shared Memory for Composed Applications on Multi-OS/R Exascale Systems
abstract
Current trends in exascale systems research indicate that heterogeneity will abound in both the hardware and software layers on future HPC systems. It is our position that exascale environments are likely to be constructed from independent partitions of hardware and system software called enclaves, with multiple enclaves co-located on the same physical nodes and each executing an optimized operating system and runtime (OS/R) to support a particular application behavior. Fully utilizing these systems will require the ability to execute composed workloads, such as in situ applications, whereby HPC simulations execute synchronously with co-located analytic packages that in turn process simulation output via shared memory. In this work, we present the design and implementation of XEMEM, a shared memory system that can efficiently construct memory mappings across enclave OSes to support composed workloads while allowing diverse application components to execute in strictly isolated enclaves. By utilizing modifications to the Kitten lightweight kernel and Palacios lightweight virtual machine monitor, as well as leveraging our recent work on lightweight "co-kernels," we demonstrate that our approach can support a diverse range of native and virtualized environments likely to be deployed on future exascale systems. Finally, we demonstrate that a multi-enclave system can reduce cross-workload contention and improve performance for a sample composed benchmark compared to a single OS approach.
Brian Kocoloski, Jack Lange
HPDC2
2015 Achieving Performance Isolation with Lightweight Co-Kernels
abstract
Performance isolation is emerging as a requirement for High Performance Computing (HPC) applications, particularly as HPC architectures turn to in situ data processing and application composition techniques to increase system throughput. These approaches require the co-location of disparate workloads on the same compute node, each with different resource and runtime requirements. In this paper we claim that these workloads cannot be effectively managed by a single Operating System/Runtime (OS/R). Therefore, we present Pisces, a system software architecture that enables the co-existence of multiple independent and fully isolated OS/Rs, or enclaves, that can be customized to address the disparate requirements of next generation HPC workloads. Each enclave consists of a specialized lightweight OS co-kernel and runtime, which is capable of independently managing partitions of dynamically assigned hardware resources. Contrary to other co-kernel approaches, in this work we consider performance isolation to be a primary requirement and present a novel co-kernel architecture to achieve this goal. We further present a set of design requirements necessary to ensure performance isolation, including: (1) elimination of cross OS dependencies, (2) internalized management of I/O, (3) limiting cross enclave communication to explicit shared memory channels, and (4) using virtualization techniques to provide missing OS features. The implementation of the Pisces co-kernel architecture is based on the Kitten Lightweight Kernel and Palacios Virtual Machine Monitor, two system software architectures designed specifically for HPC systems. Finally we will show that lightweight isolated co-kernels can provide better performance for HPC applications, and that isolated virtual machines are even capable of outperforming native environments in the presence of competing workloads.
Jiannan Ouyang, Brian Kocoloski, Jack Lange, Kevin T. Pedretti
HPDC3
2014 HPMMAP: Lightweight Memory Management for Commodity Operating Systems
abstract
Linux-based operating systems and runtimes (OS/Rs) have emerged as the environments of choice for the majority of modern HPC systems. While Linux-based OS/Rs have advantages such as extensive feature sets as well as developer familiarity, these features come at the cost of additional overhead throughout the system. In contrast to Linux, there is a substantial history of work in the HPC community focused on lightweight OS/R architectures that provide scalable and consistent performance for tightly coupled HPC applications, but lack many of the features offered by commodity OS/Rs. In this paper, we propose to bridge the gap between LWKs and commodity OS/Rs by selectively providing a lightweight memory subsystem for HPC applications in a commodity OS/R environment. Our system HPMMAP provides isolated and low overhead memory performance transparently to HPC applications by bypassing Linux's memory management layer. Our approach is dynamically configurable at runtime, and adds no additional overheads nor requires any resources when not in use. We show that HPMMAP can decrease variance and reduce application runtime by up to 50%.
Brian Kocoloski, Jack Lange
IPDPS2
2013 Virtual TCP offload: optimizing ethernet overlay performance on advanced interconnects
Patrick G. Bridges, Jack Lange, Peter A. Dinda
HPDC3
2013 Preemptable ticket spinlocks: improving consolidated performance in the cloud
abstract
When executing inside a virtual machine environment, OS level synchronization primitives are faced with significant challenges due to the scheduling behavior of the underlying virtual machine monitor. Operations that are ensured to last only a short amount of time on real hardware, are capable of taking considerably longer when running virtualized. This change in assumptions has significant impact when an OS is executing inside a critical region that is protected by a spinlock. The interaction between OS level spinlocks and VMM scheduling is known as the Lock Holder Preemption problem and has a significant impact on overall VM performance. However, with the use of ticket locks instead of generic spinlocks, virtual environments must also contend with waiters being preempted before they are able to acquire the lock. This has the effect of blocking access to a lock, even if the lock itself is available. We identify this scenario as the Lock Waiter Preemption problem. In order to solve both problems we introduce Preemptable Ticket spinlocks, a new locking primitive that is designed to enable a VM to always make forward progress by relaxing the ordering guarantees offered by ticket locks. We show that the use of Preemptable Ticket spinlocks improves VM performance by 5.32X on average, when running on a non paravirtual VMM, and by 7.91X when running on a VMM that supports a paravirtual locking interface, when executing a set of microbenchmarks as well as a realistic e-commerce benchmark.
Jiannan Ouyang, Jack Lange
VEE2
2012 A case for dual stack virtualization: consolidating HPC and commodity applications in the cloud
abstract
With the growth of Infrastructure as a Service (IaaS) cloud providers, many have begun to seriously consider cloud services as a substrate for HPC applications. While the cloud promises many benefits for the HPC community, it currently does not come without drawbacks for application performance. These performance issues are generally the result of resource contention as multiple VMs compete for the same hardware. This contention culminates in cross VM interference whereby one VM is able to impact the performance of another. For HPC applications this interference can have a dramatic impact on scalability and performance. In order to fully support HPC applications in the cloud, services need to be available that prevent cross VM interference and isolate HPC workloads from other users. As a means to achieve this goal, we propose a dual stack approach to IaaS cloud services that utilizes multiple concurrent VMMs on each node capable of partitioning local resources in order to provide performance isolation. Each partition can then be managed by a specialized VMM that is designed specifically for either an HPC or commodity environment. In this paper we demonstrate the use of the Palacios VMM, a virtual machine monitor specifically designed for HPC, in concert with KVM to provide a partitioned cloud platform that is capable of hosting both commodity and HPC applications on a single node without interference. Furthermore, our results demonstrate that running KVM and Palacios in parallel allows an HPC application to achieve isolated and scalable performance while sharing hardware resources with commodity VMs.
Brian Kocoloski, Jiannan Ouyang, Jack Lange
SoCC3
2012 Dynamic adaptive virtual core mapping to improve power, energy, and performance in multi-socket multicores
abstract
Consider a multithreaded parallel application running inside a multicore virtual machine context that is itself hosted on a multi-socket multicore physical machine. How should the VMM map virtual cores to physical cores? We compare a local mapping, which compacts virtual cores to processor sockets, and an interleaved mapping, which spreads them over the sockets. Simply choosing between these two mappings exposes clear tradeoffs between performance, energy, and power. We then describe the design, implementation, and evaluation of a system that automatically and dynamically chooses between the two mappings. The system consists of a set of efficient online VMM-based mechanisms and policies that (a) capture the relevant characteristics of memory reference behavior, (b) provide a policy and mechanism for configuring the mapping of virtual machine cores to physical cores that optimizes for power, energy, or performance, and (c) drive dynamic migrations of virtual cores among local physical cores based on the workload and the currently specified objective. Using these techniques we demonstrate that the performance of SPEC and PARSEC benchmarks can be increased by as much as 66%, energy reduced by as much as 31%, and power reduced by as much as 17%, depending on the optimization objective.
Chang Bae, Lei Xia 0001, Peter A. Dinda, Jack Lange
HPDC4
2012 VNET/P: bridging the cloud and high performance computing through fast overlay networking
abstract
It is now possible to allow VMs hosting HPC applications to seamlessly bridge distributed cloud resources and tightly-coupled supercomputing and cluster resources. However, to achieve the application performance that the tightly-coupled resources are capable of, it is important that the overlay network not introduce significant overhead relative to the native hardware, which is not the case for current user-level tools, including our own existing VNET/U system. In response, we describe the design, implementation, and evaluation of a layer 2 virtual networking system that has negligible latency and bandwidth overheads in 1--10 Gbps networks. Our system, VNET/P, is directly embedded into our publicly available Palacios virtual machine monitor (VMM). VNET/P achieves native performance on 1 Gbps Ethernet networks and very high performance on 10 Gbps Ethernet networks and InfiniBand. The NAS benchmarks generally achieve over 95% of their native performance on both 1 and 10 Gbps. These results suggest it is feasible to extend a software-based overlay network designed for computing at wide-area scales into tightly-coupled environments.
Lei Xia 0001, Jack Lange, Peter A. Dinda, Patrick G. Bridges
HPDC3
2012 Optimizing overlay-based virtual networking through optimistic interrupts and cut-through forwarding
abstract
Overlay-based virtual networking provides a powerful model for realizing virtual distributed and parallel computing systems with strong isolation, portability, and recoverability properties. However, in extremely high throughput and low latency networks, such overlays can suffer from bandwidth and latency limitations, which is of particular concern if we want to apply the model in HPC environments. Through careful study of an existing very high performance overlay-based virtual network system, we have identified two core issues limiting performance: delayed and/or excessive virtual interrupt delivery into guests, and copies between host and guest data buffers done during encapsulation. We respond with two novel optimizations: optimistic, timer-free virtual interrupt injection, and zero-copy cut-through data forwarding. These optimizations improve the latency and bandwidth of the overlay network on 10 Gbps interconnects, resulting in near-native performance for a wide range of microbenchmarks and MPI application benchmarks.
Lei Xia 0001, Patrick G. Bridges, Peter A. Dinda, Jack Lange
SC5
2011 SymCall: symbiotic virtualization through VMM-to-guest upcalls
abstract
Symbiotic virtualization is a new approach to system virtualization in which a guest OS targets the native hardware interface as in full system virtualization, but also optionally exposes a software interface that can be used by a VMM, if present, to increase performance and functionality. Neither the VMM nor the OS needs to support the symbiotic virtualization interface to function together, but if both do, both benefit. We describe the design and implementation of the SymCall symbiotic virtualization interface in our publicly available Palacios VMM for modern x86 machines. SymCall makes it possible for Palacios to make clean synchronous upcalls into a symbiotic guest, much like system calls. One use of symcalls is to allow synchronous collection of semantically rich guest data during exit handling in order to enable new VMM features. We describe the implementation of SwapBypass, a VMM service based on SymCall that reconsiders swap decisions made by a symbiotic Linux guest. Finally, we present a detailed performance evaluation of both SwapBypass and SymCall.
Jack Lange, Peter A. Dinda
VEE1
2011 Minimal-overhead virtualization of a large scale supercomputer
abstract
Virtualization has the potential to dramatically increase the usability and reliability of high performance computing (HPC) systems. However, this potential will remain unrealized unless overheads can be minimized. This is particularly challenging on large scale machines that run carefully crafted HPC OSes supporting tightly-coupled, parallel applications. In this paper, we show how careful use of hardware and VMM features enables the virtualization of a large-scale HPC system, specifically a Cray XT4 machine, with < = 5% overhead on key HPC applications, microbenchmarks, and guests at scales of up to 4096 nodes. We describe three techniques essential for achieving such low overhead: passthrough I/O, workload-sensitive selection of paging mechanisms, and carefully controlled preemption. These techniques are forms of symbiotic virtualization, an approach on which we elaborate.
Jack Lange, Kevin T. Pedretti, Peter A. Dinda, Patrick G. Bridges, Chang Bae, Philip Soltero, Alex Merritt
VEE1
2010 EmNet: Satisfying The Individual User Through Empathic Home Networks
abstract
We consider optimizing the control of the wide-area link of a home router based on the needs of individual users instead of assuming a canonical user. A careful user study clearly demonstrates that measured end-user satisfaction with a given set of home network conditions is highly variable - user perception and opinion of acceptable network performance is very different from user to user. To exploit this fact we design, implement, and evaluate a prototype system, EmNet, that incorporates direct user feedback from a simple user interface layered over existing web content. This feedback is used to dynamically configure a weighted fair queuing (WFQ) scheduler on the wide-area link. We evaluate EmNet in terms of the measured satisfaction of end-users, and in terms of the bandwidth required. We compare EmNet with an uncontrolled link (the common case today), as well as with statically configured WFQ scheduling. On average, EmNet is able to increase overall user satisfaction by 20% over the uncontrolled network and by 12% over static WFQ. EmNet does so by only increasing the average application bandwidth by 6% over the static WFQ scheduler.
Jack Lange, J. Scott Miller, Peter A. Dinda
INFOCOM1
2010 Palacios and Kitten: New high performance operating systems for scalable virtualized and native supercomputing
abstract
Palacios is a new open-source VMM under development at Northwestern University and the University of New Mexico that enables applications executing in a virtualized environment to achieve scalable high performance on large machines. Palacios functions as a modularized extension to Kitten, a high performance operating system being developed at Sandia National Laboratories to support large-scale supercomputing applications. Together, Palacios and Kitten provide a thin layer over the hardware to support full-featured virtualized environments alongside Kitten's lightweight native environment. Palacios supports existing, unmodified applications and operating systems by using the hardware virtualization technologies in recent AMD and Intel processors. Additionally, Palacios leverages Kitten's simple memory management scheme to enable low-overhead pass-through of native devices to a virtualized environment. We describe the design, implementation, and integration of Palacios and Kitten. Our benchmarks show that Palacios provides near native (within 5%), scalable performance for virtualized environments running important parallel applications. This new architecture provides an incremental path for applications to use supercomputers, running specialized lightweight host operating systems, that is not significantly performance-compromised.
Jack Lange, Kevin T. Pedretti, Trammell Hudson, Peter A. Dinda, Lei Xia 0001, Patrick G. Bridges, Andy Gocke, Steven Jaconette, Michael J. Levenhagen, Ron Brightwell
IPDPS1
2008 Experiences with Client-based Speculative Remote Display
Jack Lange, Peter A. Dinda, Samuel Rossoff
USENIX ATC1
2007 Transparent network services via a virtual traffic layer for virtual machines
abstract
We claim that network services can be transparently added to existing unmodified applications running inside virtual machine environments. Examples of these network services include protocol transformations (e.g. TCP to UDT), network connection persistence during long duration unavailability (e.g. wide area VM migration), and network flow modification (e.g. local acknowledgments and Split-TCP). To demonstrate the utility of this concept, and to enable the practical implementations of these examples and others, we have developed VTL. VTL is a framework for packet modification and creation whose purpose is to modify network traffic to and from a VM, doing so transparently to the VM and its applications. We explain how to use VTL to implement the examples mentioned above and others, such as providing anonymized connectivity for a virtual machine through the Tor anonymizing network, and creating cooperative selective wormholing services for network intrusion detection systems.
Jack Lange, Peter A. Dinda
HPDC1
2007 Vortex: Enabling Cooperative Selective Wormholing for Network Security Systems
Jack Lange, Peter A. Dinda, Fabián E. Bustamante
RAID1
2005 Automatic dynamic run-time optical network reservations
abstract
Optical networking may dramatically change high performance distributed computing. One reason is that optical networks can support provisioning dynamically configurable lightpaths, a form of circuit switching, through reservations. However, to use it (and all other network reservation mechanisms), the user or developer must modify the application. We present a system, VRESERVE, that automatically and dynamically creates network reservation requests based on the inferred network demands of running distributed and/or parallel applications with no modification to the application or operating system, and no input from the user or developer. Our execution model is a collection of virtual machines interconnected by an overlay network. The overlay network infers application demands, providing a dynamic run-time assessment of the application's topology and traffic load matrix. We then reserve lightpaths corresponding to the topology and use the overlay to forward virtual network traffic over them. We evaluate our system on the OMNInet network.
Jack Lange, Ananth I. Sundararaj, Peter A. Dinda
HPDC1